Paths, Mounts, and Environment¶
What the containers see: FormattablePath templates, mounts, environment variables, and the setup hooks that run before workers start.
FormattablePath Template System¶
FormattablePath is a powerful templating system for paths that supports runtime placeholders and environment variable expansion.
How It Works¶
FormattablePath ensures that configuration values with placeholders are always explicitly formatted before use, preventing accidental use of unformatted templates.
# Example usage in config
output:
log_dir: "$HOME/logs/{job_id}/{run_name}"
container_mounts:
"$HOME/data": "/data"
"$HOME/logs/{job_id}": "/logs"
Available Placeholders¶
| Placeholder | Type | Description | Example |
|---|---|---|---|
{job_id} |
string | SLURM job ID | "12345" |
{run_name} |
string | Job name + job ID | "my-benchmark_12345" |
{head_node_ip} |
string | IP address of head node | "10.0.0.1" |
{log_dir} |
string | Resolved log directory path | "/home/user/outputs/12345/logs" |
{model_path} |
string | Resolved model path | "/models/deepseek-r1" |
{container_image} |
string | Resolved container image path | "/containers/sglang.sqsh" |
{gpus_per_node} |
int | GPUs per node | 8 |
Environment Variable Expansion¶
FormattablePath also expands environment variables using $VAR or ${VAR} syntax:
output:
log_dir: "$HOME/outputs/{job_id}/logs"
# Expands to: /home/username/outputs/12345/logs
Common environment variables:
- $HOME - User home directory
- $USER - Username
- $SLURM_JOB_ID - SLURM job ID (also available as {job_id})
Extra Placeholders¶
Some contexts support additional placeholders:
| Placeholder | Context | Description |
|---|---|---|
{nginx_url} |
Frontend config | Nginx URL for load balancing |
{frontend_url} |
Frontend config | Frontend/router URL |
{index} |
Worker config | Worker index |
{host} |
Worker config | Worker host |
{port} |
Worker config | Worker port |
Examples¶
# Log directory with job ID
output:
log_dir: "./outputs/{job_id}/logs"
# Mount user data into container
container_mounts:
"$HOME/datasets": "/datasets"
"./outputs/{job_id}": "/outputs"
# Custom paths with environment variables
extra_mount:
- "$SCRATCH/cache:/cache"
- "${DATA_DIR}/models:/models:ro"
container_mounts¶
Custom container mount mappings with FormattablePath support.
container_mounts:
"$HOME/datasets": "/datasets"
"$HOME/outputs/{job_id}": "/outputs"
"/shared/cache": "/cache"
Both keys and values support FormattablePath templating with placeholders and environment variables.
Default Mounts¶
The following mounts are always added automatically:
| Host Path | Container Path | Description |
|---|---|---|
| Model path | /model |
Resolved model directory |
| Log directory | /logs |
Log output directory |
configs/ directory |
/configs |
NATS, etcd binaries |
| Benchmark scripts | /srtctl-benchmarks |
Bundled benchmark scripts |
Cluster-Level Mounts¶
You can also define cluster-wide mounts in srtslurm.yaml using the default_mounts field. These are applied to all jobs on the cluster, after the built-in defaults but before job-level mounts.
# In srtslurm.yaml
default_mounts:
"/cluster/special/libs": "/opt/libs"
"$SCRATCH": "/scratch"
Environment variables (e.g., $SCRATCH, $HOME) are expanded. This is useful for mounting cluster-specific paths that are required by certain images without adding them to every job config.
Mount Priority¶
Mounts have the following priority (highest to lowest):
- Job-level
container_mounts- FormattablePath dict (highest priority) - Job-level
extra_mount- simplehost:containerstrings - Cluster-level -
default_mountsfromsrtslurm.yaml - Built-in defaults - model, logs, configs, benchmark scripts (lowest priority)
Job-level mounts always take precedence over cluster-level and built-in defaults.
environment¶
Global environment variables for all worker processes.
environment:
MY_VAR: "value"
CUDA_LAUNCH_BLOCKING: "1"
NCCL_DEBUG: "INFO"
Per-Worker Template Variables¶
Environment variable values support per-worker templating with these placeholders:
| Placeholder | Description | Example |
|---|---|---|
{node} |
Hostname of the node where the worker runs | "gpu-01" |
{node_id} |
Numeric index of the node in worker list (0-based) | 0, 1, 2 |
Note: For per-role environment variables, use roles.prefill.env, roles.decode.env, or roles.agg.env (see roles). The role's env is applied first and the global environment after it, so a key set in both takes the global value.
extra_mount¶
Additional container mounts as a list of mount specifications.
extra_mount:
- "/local/path:/container/path"
- "/data:/data:ro"
- "$HOME/cache:/cache"
| Format | Description |
|---|---|
host_path:container_path |
Read-write mount |
host_path:container_path:ro |
Read-only mount |
Note: Unlike container_mounts, extra_mount uses simple string format, not FormattablePath. Environment variables are still expanded.
setup_script¶
Run a custom script before dynamo install and worker startup.
setup_script: "install-custom-deps.sh"
Notes:
- Script must be located in the
configs/directory. - Script runs inside the container before dynamo installation.
- Useful for installing custom SGLang versions, additional dependencies, or patches.
Example setup script (configs/install-sglang-main.sh):
#!/bin/bash
pip install --quiet git+https://github.com/sgl-project/sglang.git
host_setup¶
Commands run on each allocated node's bare host, outside the container, before any worker starts.
This is the counterpart to setup_script, which runs inside the container. Use host_setup for node state the container cannot reach: locking GPU clocks, loading a kernel module, dropping caches.
host_setup:
commands:
- "sudo -n nvidia-smi -lmc <min>,<max>"
teardown:
- "sudo -n nvidia-smi -rmc"
nodes: all
ignore_failure: false
timeout_seconds: 300
Fields: HostSetupConfig.
How it runs: the orchestrator itself runs on the host (not in a container), so it fans these out as one container-less srun per node, in parallel. Output lands in <log_dir>/host_setup_<node>.out and <log_dir>/host_teardown_<node>.out.
Notes:
- Commands run as you, not as root. Anything privileged needs passwordless sudo (
sudo -n ...). Asudothat prompts for a password will hang untiltimeout_secondsand then fail the job; verify first withsrun --jobid <job> --overlap -w <node> sudo -n true. If sudo prompts, no recipe change helps; the cluster's SLURMProlog=(which runs as root) is the only route. - Prefer setting
teardownwhenevercommandschanges persistent node state.nvidia-smi -lmcoutlives the allocation, so without a matching-rmcthe next job on that node inherits your locked clocks.srtctl dry-runwarns whencommandsis set withoutteardown. teardownruns from the job's cleanup path, so it fires on failure and cancellation too, and never changes the job's exit code.- Set cluster-wide via
default_host_setupinsrtslurm.yaml; that's the right home when the cluster's machines need this, rather than one recipe. See Cluster Config Fields. srtctl dry-run -f config.yamlrenders the commands, their scope, and which file they came from.- For more than a couple of commands, or when you want each one logged with its exit code, point both lists at
configs/node-hooks.shand write the commands one per line inHOOK_PRE/HOOK_POSTblock scalars underenvironment:. The job script exports those, so the host sruns inherit them. See examples/features/node-hooks.yaml.