UV / Pip¶
Python >= 3.12. Install TServe from PyPI with uv or pip, then start tserve. Editable installs from a clone stay on From source.
Install¶
The gpu extra does not change a pip install. Family extras already install CUDA torch from PyPI (MPS on macOS). To force a CPU wheel, install torch separately first — CPU-only install.
server is enough to serve naive (a test baseline). The hub extra above covers Chronos Bolt/T5, TTM, and TimesFM 2.x. Do not add client on a machine that only serves.
Dependencies¶
The extra name is the CPU image tag. GPU tags are {extra}-gpu. There is no :base-gpu.
kronos is built on base, so it cannot load Chronos Bolt, TTM, or TimesFM. Every extra that pulls hf — hub, chronos, granite, moirai, tirex, tirex2, toto, mantis, timesfm3, t0, tafsut, and full — can.
full is chronos, kronos, granite, moirai, tirex, tirex2, toto, mantis, timesfm3, t0, and tafsut. client, http, gpu, dev, docs, and all-extras are not model families. gpu selects the torch index for uv on a clone only. Pip ignores gpu and installs CUDA torch from PyPI (MPS on macOS). Moirai pins gluonts, lightning, and hydra-core when python_version < '3.14'.
added counts checkpoints that extra contributes. full is the total, including naive.
| extra | CPU tag | GPU tag | families | added | example |
|---|---|---|---|---|---|
server |
:base |
— | Naive | 1 | naive |
hub |
:hub |
:hub-gpu |
Chronos Bolt, Chronos T5, TTM, TimesFM 2.x | 81 | chronos_bolt |
chronos |
:chronos |
:chronos-gpu |
Chronos-2 | 3 | chronos_2 |
kronos |
:kronos |
:kronos-gpu |
Kronos, WindFM | 5 | kronos |
granite |
:granite |
:granite-gpu |
FlowState | 2 | flowstate |
moirai |
:moirai |
:moirai-gpu |
Moirai 2, Moirai 1.x, Lag-Llama | 8 | moirai_2 |
tirex |
:tirex |
:tirex-gpu |
TiRex | 2 | tirex |
tirex2 |
:tirex2 |
:tirex2-gpu |
TiRex-2 | 4 | tirex_2 |
toto |
:toto |
:toto-gpu |
Toto-2 | 5 | toto_2_0_4m |
mantis |
:mantis |
:mantis-gpu |
Mantis | 3 | mantis_8m |
timesfm3 |
:timesfm3 |
:timesfm3-gpu |
TimesFM 3 | 1 | timesfm_3 |
t0 |
:t0 |
:t0-gpu |
T0 | 1 | t0 |
tafsut |
:tafsut |
:tafsut-gpu |
Tafsut | 1 | tafsut |
full |
:full |
:full-gpu |
all of the above | 117 | chronos_2 |
Replace hub in the install above with another extra from the table. full is the union extra. all-extras is a pip convenience for client,server,full and is not a Docker tag. Each extra's command and models: catalog.
gpu is not a model family. A CPU wheel: CPU-only install.
CPU-only install¶
Family extras pull torch, and every install here takes the CUDA wheel from PyPI (MPS on macOS). On a GPU host that is already what you want, so nothing below is needed.
Without a GPU, that wheel is a large download you will never use. Install torch from the CPU index first, then TServe. If a later install replaces that wheel with CUDA, run the torch line again.
The same order works with uv pip install. Swap hub for any other family extra from Dependencies.
The gpu extra only selects the torch index in a clone's uv lockfile: From source. Containers pick the wheel through the tag instead: Docker.
Serve from the command line¶
Startup prints the URLs it binds:
Starting TServe
Dashboard http://127.0.0.1:8000/
Swagger UI http://127.0.0.1:8000/docs
ReDoc http://127.0.0.1:8000/redoc
--host 127.0.0.1 accepts local connections only. 0.0.0.0 also accepts them from your network. --log-level is debug, info, warning, error, or critical. Ctrl+C stops the process and exits 0. Every flag: CLI. A craft spec is id=spec: Craft specs.
Serve from Python¶
Server takes the same arguments:
from tserve.server import Server
server = Server(
model=["chronos_bolt", "ttm_r3"],
host="127.0.0.1",
port=8000,
)
print(server.url) # http://127.0.0.1:8000
server.run()
Models load during construction, so an unknown model or a missing dependency raises before the port is bound. run() blocks until the process stops. Omitting model still loads naive.
server.app is the FastAPI app:
Use one worker per process. Each worker loads its own copy of every model.
Next¶
- From source — editable install from a clone
- Live objects — Python can also serve estimators you configured in the session
- Craft specs — load a sktime craft spec as
(id, spec)or CLIid=spec - Models from a directory — serve saved sktime
.zipfiles - Dashboard — the console at http://127.0.0.1:8000/