Skip to content

Recipe Guide

How to write a job recipe in the 2.0 (schema: 2) layout: what each block means, how the blocks interact, and worked examples.

Field-level facts (every key, type, default, and allowed value) are generated from the code into schema-reference.md and checked in CI; if a prose page and the generated one disagree, the generated one is right. Agents and editors can read the same data as JSON Schema (srtctl schema, published at schema/recipe.schema.json) or per field from the MCP explain_field tool. The pre-2.0 layout (backend:, resources.prefill_nodes, infra:, ...) is in legacy-v1.md; srtctl migrate -f recipe.yaml --in-place rewrites it.

Topic Page Recipe keys
Cluster defaults and aliases Cluster Config srtslurm.yaml
Engine and engine-wide knobs Engines engine
Nodes, GPUs, Slurm Topology and Placement roles, placement, resources, slurm, sbatch_directives, srun_options
Router and Dynamo Frontends and Dynamo frontend, dynamo
Load and evals Benchmarks benchmark, post_eval
Profilers Profiling profiling
Metrics and power Observability and Telemetry observability, telemetry
Container paths and env Paths, Mounts, and Environment container_mounts, extra_mount, environment, setup_script, host_setup
Sidecars and infrastructure Services services
Many jobs from one file Parameter Sweeps, Config Overrides sweep, base / override_*

This page covers the remaining top-level keys (schema, name, model, output, health_check, enable_config_dump) and complete examples.

schema

schema: 2 is the layout this document describes and the only one that loads. A recipe without the key is the pre-2.0 layout in legacy-v1.md and is rejected.

Put the key first in the file, beside base: in an override file. Upgrade a recipe with srtctl migrate -f recipe.yaml --in-place, which preserves comments and key order and folds the legacy layout into engine:, roles:, placement:, services:, and dynamo.source (a directory is walked recursively). A schema: 2 recipe that still carries a pre-2.0 key (backend:, infra:, resources.prefill_nodes, dynamo.hash, ...) is rejected too; the error names the keys and srtctl migrate rewrites them.

schema: 2
name: "deepseek-r1-benchmark"

Two rules worth knowing: a benchmark: field the selected type does not read is a load error (see benchmark), and roles.<role>.nodes: 0 is rejected in favor of the explicit colocate (see roles).

name

Job name: the Slurm --job-name (unless RUNNER_NAME is set) and the run's label in results.

name: "deepseek-r1-benchmark"

model

Model and container configuration.

model:
  path: "deepseek-r1"       # Alias from srtslurm.yaml or full path
  container: "sglang"       # Container alias from srtslurm.yaml
  precision: "fp8"          # fp8, fp4, bf16, etc.

Fields: ModelConfig.

output

Output configuration with formattable paths.

output:
  log_dir: "./outputs/{job_id}/logs"

The log_dir supports FormattablePath templating. See FormattablePath Template System.

health_check

Health check configuration for worker readiness, and the worker log watch that fails a run early.

health_check:
  max_attempts: 180
  interval_seconds: 10
  fatal_log_markers: true
  extra_fatal_log_patterns: []

Fields: HealthCheckConfig.

Notes:

  • Default of 180 attempts at 10 second intervals = 30 minutes total wait time.
  • Large models (e.g., 70B+ parameters) may require the full 30 minutes to load.
  • Reduce max_attempts for smaller models or faster testing.

Worker log watch (fatal_log_markers): the process monitor normally learns that a worker died from its srun step exiting. A TRT-LLM worker step is one trtllm-llmapi-launch task per GPU and the engine is a child of the rank-0 task only; when that child dies the launcher prints Rank0 Task exit code: <n> and the other ranks stay blocked, so the step never exits and the run would otherwise wait out the whole health window. With the watch on, the monitor scans each critical worker's log for the lines its engine declares fatal (TRT-LLM: Rank<N> Task exit code: <non-zero> and Failed to initialize executor; other engines declare none) and fails the run within one monitor poll, printing the process name and the matching line. Bare words such as Traceback or MPI_Abort are deliberately not markers: Dynamo logs a traceback for every request cancelled at EOS, and MPI abort lines appear on normal teardown. Use extra_fatal_log_patterns to add a marker for one recipe (for example "CUDA error: out of memory"), and fatal_log_markers: false to switch the watch off, for probes that kill workers on purpose. TRT-LLM endpoint steps are also launched with srun --kill-on-bad-exit=1, so a launcher task that does exit non-zero ends the whole step.

enable_config_dump

Accepted for compatibility; srtctl does not read it.

enable_config_dump: true

SGLang (except under the sglang direct frontend) and vLLM workers always get --dump-config-to, which writes the resolved engine configuration to a JSON file in the log directory.

Complete Examples

Every example below loads with srtctl dry-run. Model and container names are srtslurm.yaml aliases.

Disaggregated Mode with Dynamo

schema: 2
name: "deepseek-r1-disagg"

model:
  path: "deepseek-r1"
  container: "sglang"
  precision: "fp8"

resources:
  gpu_type: "gb200"
  gpus_per_node: 4

slurm:
  time_limit: "04:00:00"

dynamo:
  source:
    pypi: "1.4.2"

frontend:
  type: dynamo
  enable_multiple_frontends: true
  args:
    router-mode: "kv"

engine: sglang
roles:
  prefill:
    nodes: 2
    workers: 4
    gpus: 2
    kv_events: true
    env:
      TORCH_DISTRIBUTED_DEFAULT_TIMEOUT: "1800"
    args:
      tensor-parallel-size: 2
      mem-fraction-static: 0.84
      kv-cache-dtype: "fp8_e4m3"
  decode:
    nodes: 4
    workers: 2
    gpus: 8
    env:
      TORCH_DISTRIBUTED_DEFAULT_TIMEOUT: "1800"
    args:
      tensor-parallel-size: 8
      mem-fraction-static: 0.83
      data-parallel-size: 8

benchmark:
  type: "sa-bench"
  isl: 1024
  osl: 1024
  concurrencies: [128, 256, 512]

health_check:
  max_attempts: 180
  interval_seconds: 10

Aggregated Mode with SGLang Router

schema: 2
name: "qwen-agg-router"

model:
  path: "qwen3-32b"
  container: "sglang"
  precision: "bf16"

resources:
  gpu_type: "h100"
  gpus_per_node: 8

slurm:
  time_limit: "02:00:00"

frontend:
  type: sglang-router
  enable_multiple_frontends: false
  args:
    policy: "cache_aware"

engine: sglang
roles:
  agg:
    nodes: 4
    workers: 8
    gpus: 4
    args:
      tensor-parallel-size: 4
      mem-fraction-static: 0.9
      enable-dp-attention: true

benchmark:
  type: "router"
  isl: 14000
  osl: 200
  num_requests: 200
  prefix_ratios: [0.1, 0.3, 0.5, 0.7, 0.9]

Profiling Example

schema: 2
name: "profile-decode"

model:
  path: "llama-70b"
  container: "sglang"
  precision: "fp8"

resources:
  gpu_type: "h100"
  gpus_per_node: 8

slurm:
  time_limit: "01:00:00"

profiling:
  type: "torch"
  prefill:
    start_step: 5
    stop_step: 15
  decode:
    start_step: 5
    stop_step: 15

engine: sglang
roles:
  prefill:
    nodes: 1
    workers: 1
    gpus: 8
    args:
      tensor-parallel-size: 8
  decode:
    nodes: 1
    workers: 1
    gpus: 8
    args:
      tensor-parallel-size: 8

benchmark:
  type: "sa-bench"
  isl: 2048
  osl: 256
  concurrencies: "32x64"
  req_rate: "inf"

Parameter Sweep Example

schema: 2
name: "sweep-throughput"

model:
  path: "deepseek-r1"
  container: "sglang"
  precision: "fp8"

resources:
  gpu_type: "gb200"
  gpus_per_node: 4

engine: sglang
roles:
  prefill:
    nodes: 1
    workers: 2
    gpus: 2
    args:
      tensor-parallel-size: 2
  decode:
    nodes: 2
    workers: 4
    gpus: 2
    args:
      tensor-parallel-size: 2

benchmark:
  type: "sa-bench"
  isl: "{isl}"
  osl: "{osl}"
  concurrencies: [64, 128, 256]

sweep:
  isl: [512, 1024, 2048, 4096]
  osl: [128, 256, 512, 1024]

Config Override Example

schema: 2
base:
  name: "disagg-fp8-benchmark"

  model:
    path: "deepseek-r1"
    container: "sglang"
    precision: "fp8"

  resources:
    gpu_type: "h100"
    gpus_per_node: 8

  engine: sglang
  roles:
    prefill:
      nodes: 2
      workers: 2
      gpus: 8
      args:
        tp-size: 8
    decode:
      nodes: 8
      workers: 8
      gpus: 8
      args:
        tp-size: 8

  benchmark:
    type: "sa-bench"
    isl: 1024
    osl: 8192
    concurrencies: [8192, 10240]

# One TP=16 decode worker spanning two nodes, prefill unchanged
override_tp16:
  roles:
    decode:
      workers: 4
      gpus: 16
      args:
        tp-size: 16

# Smaller cluster with fewer decode nodes
override_small:
  roles:
    decode:
      nodes: 4
      workers: 4
  benchmark:
    concurrencies: [4096]

Custom Mounts and Setup

schema: 2
name: "custom-setup"

model:
  path: "$MODELS_DIR/my-model"
  container: "$CONTAINERS_DIR/custom.sqsh"
  precision: "fp8"

resources:
  gpu_type: "h100"
  gpus_per_node: 8

engine: sglang
roles:
  agg:
    nodes: 2
    workers: 4
    gpus: 4
    args:
      tensor-parallel-size: 4

setup_script: "install-custom-sglang.sh"

environment:
  CUSTOM_VAR: "value"
  NCCL_DEBUG: "INFO"

container_mounts:
  "$HOME/datasets": "/datasets"
  "$SCRATCH/cache": "/cache"

extra_mount:
  - "/shared/data:/data:ro"

sbatch_directives:
  mail-user: "user@example.com"
  mail-type: "END,FAIL"
  reservation: "gpu-cluster"

srun_options:
  cpu-bind: "none"

output:
  log_dir: "$HOME/experiments/{job_id}/logs"

health_check:
  max_attempts: 120
  interval_seconds: 15