Engines¶
The top-level engine: block: which inference engine runs, and the engine-wide knobs whose behavior needs more than a table row. Per-engine field tables: Engine types (generated).
engine¶
GPU scheduling uses upstream's existing cluster settings. For eight-GPU
allocations on GRES-only clusters, set use_gpus_per_node_directive: false
and default_sbatch_directives: {gres: "gpu:8"}.
GPU visibility on AMD¶
Set visible_devices_env: ROCR_VISIBLE_DEVICES in the cluster profile for ROCm
workers. GPU subsets then use only that mask, without applying a second mask to
already-renumbered devices. Set default_gpu_exporter: null to disable the
NVIDIA GPU exporter, or configure an exporter image, port, and command once for
the cluster. Other telemetry is unchanged; an explicit recipe exporter wins.
For vLLM builds without --device-ids, set engine.set_visible_devices: true.
This is one explicit boolean, not automatic vLLM version detection. The default
is false: vLLM binds devices with --device-ids. There is no CUDA-named alias.
engine: supplies the default inference engine for worker roles. It is optional when every role sets its own engine. A bare string is the common form; a mapping carries engine options:
engine: sglang
engine:
type: vllm
connector: nixl # vLLM KV connector for disaggregation
engine:
type: trtllm
served_model_name: "Qwen/Qwen3-0.6B"
For trtllm_serve, an explicit served_model_name is passed to the worker's
--served_model_name option so the server and benchmark/eval clients use the same
API model name. Leave it unset to retain the server's default. This is a server
option, not a key in roles.*.args (the engine YAML).
engine:
type: mocker
engine_type: vllm
speedup_ratio: 100
Valid types are atom, sglang, tilert, trtllm, vllm, and mocker. Everything that is per role (the role's environment, its engine CLI flags, extra_args, kv_events) lives under roles; everything else about the engine lives here. The generated tables under Engine types list every engine-wide knob per engine; the ones worth knowing are:
| Engine | Engine-wide knobs |
|---|---|
sglang-router |
none beyond type |
vllm |
connector (default nixl), dp_launch_mode, vllm_serve_binary, set_visible_devices, allow_prefill_decode_colocation, allow_prefill_decode_colocation_across_nodes |
trtllm |
served_model_name, publish_metrics, publish_events_and_metrics, sequential_node_start, numa_memory_bind, numa_cpu_bind |
mocker |
the simulation parameters: engine_type, speedup_ratio, decode_speedup_ratio, num_gpu_blocks_override, max_num_seqs, max_num_batched_tokens, block_size, data_parallel_size, ... |
The v1 spelling of this (backend.type plus the engine-wide keys under backend:) is documented in legacy-v1.md; srtctl migrate rewrites it.
TRT-LLM CPU and memory placement¶
To place worker CPUs and memory on the NUMA node associated with each task's GPU:
engine:
type: trtllm
numa_cpu_bind: true
numa_memory_bind: local
The launcher resolves the GPU through CUDA_VISIBLE_DEVICES and
SLURM_LOCALID, applies its CPU mask, and sets numactl --membind=<node>
before starting the worker. Allocations governed by this policy cannot fall
back to another node. Insufficient local memory can cause allocation failure
or OOM, even when another node has free memory. Existing or shared pages are
not migrated. The container must provide numactl. Local mode requires
numa_cpu_bind: true. The wrapper uses CUDA_VISIBLE_DEVICES; alternate
cluster GPU visibility variables are not supported by this wrapper.
numa_memory_bind: false keeps CPU binding without a memory policy change.
numa_memory_bind: true uses numactl -m 0,1 for any GPU type or worker mode.
When omitted or null, this two-node policy applies only to gb200, gb300,
and vrnvl72 prefill and decode workers. Enabling CPU binding does not change
these memory policies.
In local mode, the launcher fails if the GPU's NUMA affinity cannot be resolved,
its CPU list is missing or empty, or the memory policy cannot be applied.
Without local mode, unknown GPU NUMA affinity skips CPU binding and retains
the selected memory policy. When profiling in local mode, the outer nsys
process also inherits the strict memory policy.
See the local-binding example.
vLLM DP launch mode¶
vLLM data-parallel endpoints use one process per node by default. srtslurm derives whether each TP/PP replica is node-local or spans multiple nodes:
engine: vllm
roles:
prefill:
args:
data-parallel-size: 8
decode:
args:
data-parallel-size: 16
| Value | Process layout |
|---|---|
per_node |
One process per node (default); supports node-local or distributed TP/PP |
per_gpu |
One process per DP rank (TP x PP GPUs each; deprecated compatibility mode) |
Set engine.dp_launch_mode: per_gpu only when temporarily preserving the legacy process layout. srtslurm emits a configuration-time deprecation warning for Dynamo-backed DP configurations that select it. per_gpu will be removed in a future release.
When TP x PP fits on one node, srtslurm derives --data-parallel-size-local and --data-parallel-start-rank, then enables --data-parallel-hybrid-lb so every node-local process registers with the Dynamo frontend. When TP x PP is larger than the node-local GPU allocation, srtslurm instead derives the multi-node rendezvous arguments and makes every process except the global leader headless. For example, both DP4 x TP4 and DP2 x TP8 are selected automatically on four-GPU nodes.
Do not set data-parallel-size-local, data-parallel-start-rank, data-parallel-hybrid-lb, or headless manually; srtslurm owns those values. The allocation must be regular: DP x TP x PP must match the endpoint GPU count, and a TP/PP replica must divide evenly within or across nodes.
TRT-LLM metrics publication¶
With frontend.type: dynamo, prefill, decode, and aggregated TRT-LLM workers publish engine metrics by default using --publish-metrics, regardless of whether observability is enabled. publish_events_and_metrics is retained only for backward compatibility with older Dynamo builds. Observability does not enable this legacy flag automatically.
engine:
type: trtllm
publish_metrics: true # default: metrics only
publish_events_and_metrics: false # use publish_metrics
| Field | Type | Default | Description |
|---|---|---|---|
publish_metrics |
bool | true | Pass --publish-metrics unless the legacy combined flag is selected |
publish_events_and_metrics |
bool or null | unset | true: pass only the legacy combined flag; false or unset/null: use publish_metrics |
The flags are mutually exclusive. Explicit engine.publish_events_and_metrics: true passes only --publish-events-and-metrics, even when publish_metrics is true. False, omitted, and null all select the metrics-only setting.
publish_events_and_metrics |
publish_metrics |
Publication flag (with or without observability) |
|---|---|---|
omitted / null / false |
true |
--publish-metrics |
omitted / null / false |
false |
none |
true |
either | --publish-events-and-metrics |
Compatibility: the metrics-only flag requires a Dynamo build containing ai-dynamo/dynamo#12162 or equivalent support. For older builds, set engine.publish_events_and_metrics: true to select the legacy flag, which also enables KV events. To omit both flags, set engine.publish_metrics: false and leave the legacy flag false or unset. Omitting the flag does not override metrics-related environment variables supplied by the user. Metrics collection adds engine telemetry work; metrics-only does not mean zero overhead, but with the iteration-statistics default below the remaining cost is the per-request perf metrics.
Migration from the previous publication behavior: engine.publish_events_and_metrics: false previously omitted both publication flags. It now uses publish_metrics, which defaults to true. Recipes that used false to disable publication, especially on older Dynamo builds that reject --publish-metrics, must also set engine.publish_metrics: false. To publish metrics on those older builds, select engine.publish_events_and_metrics: true instead. Null remains accepted for existing serialized recipes and behaves the same as false.
KV events and observability: observability.enabled: true no longer automatically enables TRT-LLM KV events. Router KV-event dashboard panels (ro_kv_events_applied, ro_kv_event_warnings, and ro_kv_events_dropped) require an explicit event-publication opt-in. On Dynamo builds supporting the independent controls, keep the default metrics-only flag and set DYN_TRTLLM_PUBLISH_KV_EVENTS: "true" in every worker role that should publish events:
engine:
type: trtllm
roles:
agg:
env:
DYN_TRTLLM_PUBLISH_KV_EVENTS: "true"
For disaggregated recipes, set the same environment variable under both roles.prefill.env and roles.decode.env. This uses Dynamo's independent KV-event control; it does not require the legacy combined flag. Older builds without that control must use engine.publish_events_and_metrics: true to enable events and metrics together.
Iteration statistics default. srtctl bakes enable_iter_perf_stats: false into every TRT-LLM engine section a recipe uses (prefill and decode, or aggregated), under both frontend.type: dynamo and trtllm_serve, creating the section when the recipe has none. This is a setdefault: an explicit enable_iter_perf_stats: true in the recipe wins, and observability.enabled: true keeps its own true because its expansion runs first. The default exists because dynamo.trtllm turns --publish-metrics into enable_iter_perf_stats: true in the engine arguments, and the engine YAML is merged over those arguments and wins on conflicts; without the explicit key every Dynamo worker collects TensorRT-LLM's per-iteration statistics (KV-cache stats and CUDA-event step timing on every executor loop). The request-level trtllm_* series (request latency, TTFT, TPOT, queue/prefill/decode time, token counters) do not need the key: they come from the per-request perf metrics, which --publish-metrics sets on the Dynamo path and return_perf_metrics: true sets for trtllm-serve. What the default drops is the iteration-level trtllm_* gauges (trtllm_kv_cache_*, running/waiting requests, iteration latency) and, on Dynamo, the dynamo_component_kvstats_* gauges, the router worker-load sample and the Planner's forward-pass metrics; set the key to true or enable observability to get them back. One visible effect to expect on a default run: the component dashboard's engine-tab KV-cache utilisation and hit-rate panels have no data, and the Dynamo bench dashboard's KV-utilisation series sits at the gauge's seeded 0 %, because both read gauges that only iteration statistics update. Engine sections whose backend is the legacy tensorrt engine are left alone: its LlmArgs rejects the key on containers older than TensorRT-LLM v1.3.0rc21, and that backend always collected the statistics anyway.
roles:
decode:
args:
enable_iter_perf_stats: true # opt back in for one role
Saved and locked recipes carry the resolved key, so a later observability.enabled: true on such a file meets an explicit false rather than an omission; srtctl warns at load time and srtctl dry-run shows the value, and removing the line restores the default.
These options do not change native trtllm_serve or sidecar worker commands. srtctl dry-run shows the publication flag selected for Dynamo TRT-LLM workers and, for every TRT-LLM backend, the per-role enable_iter_perf_stats / return_perf_metrics values the engine YAML will carry.
TRT-LLM workers can span partially occupied nodes. For example, on four-GPU nodes, two DEP6 prefill workers need three nodes:
roles:
prefill:
nodes: 3
workers: 2
gpus: 6
args:
tensor_parallel_size: 6
moe_expert_parallel_size: 6
pipeline_parallel_size: 1
enable_attention_dp: true
The workers use A[0,1,2,3] + B[0,1] and C[0,1,2,3] + B[2,3].
Each endpoint launches exactly six MPI ranks and gets its own per-node
CUDA_VISIBLE_DEVICES. Full nodes lead the rank order to keep TRT-LLM's local
device mapping consistent. Layouts incompatible with that mapping (for example,
seven ranks split 4+3) are rejected before launch. Backend-specific communication
requirements still apply; this does not enable arbitrary uneven layouts in every
TRT-LLM communication backend.
TRT-LLM endpoints also set MASTER_ADDR to the rank-zero node and use a distinct
MASTER_PORT per endpoint. This overrides container hooks that infer rank zero
from Slurm's sorted node list. Explicit recipe environment values take precedence.
Other TRT-LLM launch facts: TRT-LLM supports prefill, decode, and aggregated roles, uses MPI-style launching (one srun per endpoint with all of its nodes) through trtllm-llmapi-launch, and sets TRTLLM_EPLB_SHM_NAME to a unique UUID per endpoint.
ATOM with AToMesh¶
Use engine: atom with frontend.type: atomesh to launch native
atom.entrypoints.openai_server workers and the official AToMesh router. Both
aggregate workers and prefill/decode topologies use static HTTP endpoints;
disaggregated workers receive topology-owned Mooncake handshake ports.
Engine flags belong under roles.prefill.args, roles.decode.args, or
roles.agg.args (schema v2).
srt-slurm owns the model path, HTTP port, tensor parallel size, and KV-transfer
contract, so recipes cannot override those arguments. See the complete
ATOM/AToMesh recipe.
To run another KV connector next to the Mooncake one (for example ATOM's in-process
LMCache CPU offload on prefill), list it under extra-kv-connectors in that role's
args. srtctl keeps generating the Mooncake entry, with its handshake port, and wraps
both in ATOM's multi connector. An aggregate role with one extra connector runs it
on its own. An lmcache_mp connector (ATOM's client for the lmcache-server
service) dials the server on its own node, port 8750, unless its
kv_connector_extra_config sets lmcache.mp.port or lmcache.mp.server_urls.
roles:
prefill:
env:
PYTHONHASHSEED: "0" # LMCache prefix hashes must agree across offload workers
args:
extra-kv-connectors:
- kv_connector: lmcache_offload
kv_role: offload
lmcache.local_cpu: true
lmcache.max_local_cpu_size: 180 # GB per worker
lmcache.chunk_size: 256
schema: 2 # Required: recipe layout version
name: "my-benchmark" # Required: job name
model: # Required: model settings
path: "deepseek-r1"
container: "sglang"
precision: "fp8"
resources: # Cluster facts: GPU type and GPUs per node
gpu_type: "gb200"
gpus_per_node: 4
slurm: # Optional: SLURM overrides
time_limit: "02:00:00"
frontend: # Optional: router/frontend config
type: dynamo
engine: sglang # Default engine; omit when every role sets engine
roles: # Required: one block per worker role
prefill:
nodes: 1
workers: 2
gpus: 2
args:
tensor-parallel-size: 2
decode:
nodes: 2
workers: 2
gpus: 4
args:
tensor-parallel-size: 4
benchmark: # Optional: benchmark config
type: "sa-bench"
isl: 1024
osl: 1024
concurrencies: [256, 512]
dynamo: # Optional: where Dynamo comes from
source:
pypi: "1.4.2"
profiling: # Optional: profiling config
type: "none"
output: # Optional: output paths
log_dir: "./outputs/{job_id}/logs"
health_check: # Optional: health check settings
max_attempts: 180
interval_seconds: 10
setup_script: "my-setup.sh" # Optional: custom setup script
The v1 spelling of this layout (backend:, resources.prefill_nodes and friends, infra:, dynamo.version) is documented in legacy-v1.md; srtctl migrate rewrites it.