Skip to content

TileRT decode

Use roles.prefill.engine: vllm, roles.decode.engine: tilert, and frontend.type: tilert-router. Do not set a top-level engine alongside role engines. Set frontend.enable_multiple_frontends: false. Recipes:

The supported pairing is vLLM prefill + TileRT decode + TileRT router. Configuration loading rejects other engine/router combinations before Slurm submission. Other prefill engines need a TileRT protocol integration and an update to the frontend adapter's validation; using NIXL or Mooncake alone is not enough. Set the model profile, sequence limit, KV layout, and transport consistently in vLLM's TileRTConnector and the decode arguments. The configuration checks do not verify the connector installed inside the image or prove KV-layout compatibility.

The adapter follows TileRT 0.1.6's decode server and router.

Images and weights

Use images containing the engine and required connector. Set roles.prefill.container and roles.decode.container for the worker images. model.container supplies the shared image for non-worker tasks. The router image defaults to model.container; use frontend.container_image to override it. It must contain tilert.pd_vllm.pd_router and its dependencies. The B200 recipe's image aliases must be defined in srtslurm.yaml.

Convert weights with TileRT before serving and mount them at roles.decode.args.model-weights-dir. Set frontend.args.model-path to the tokenizer path or Hugging Face model ID, including when parser: none.

Limits

The router supports one prefill worker and one or more decode workers, each decode worker on a single node. TileRT requires transferred KV state and cannot serve aggregate requests through this adapter.

srtctl sets the decode HTTP/control ports and --engine tilert; other options come from roles.decode.args. Workers must pass /health before the router starts. TileRT decode has no /metrics, so only prefill metrics are scraped.

The role-engine restrictions also apply.

Benchmarks

Use custom with a client that sends /v1/chat/completions and an explicit model name, as in the InferenceX validations. The public model name comes from vLLM prefill's served-model-name; it need not be repeated on the decode engine. An explicit, conflicting decode model name is rejected.

TileRT 0.1.6 has no /v1/models endpoint and only supports streaming on /v1/chat/completions. The built-in sa-bench runner uses streaming /v1/completions and is therefore incompatible. GPQA's automatic model discovery also cannot work; lm-eval needs an explicit MODEL_NAME in benchmark.env to bypass discovery. Neither evaluation path was validated by the TileRT sweeps.