SGLang Fast Engine Recovery (weight cache daemon)¶
SGLang's weight cache daemon is a persistent GPU process that loads the model once, keeps the post-quantized, TP-sharded tensors in HBM, and hands CUDA IPC handles to any engine on the same GPU. An engine started with --weight-cache-mode client maps those tensors instead of reading the checkpoint. The rest of engine startup (CUDA graphs, kernel JIT, tokenizer) still runs. The flags and limitations below follow SGLang v0.5.20's implementation.
srtctl needs no code for this. The daemon is a service with placement.per: worker, the engine flag is an ordinary roles.<role>.args entry, and roles.<role>.restart is what relaunches a dead engine so the fast load pays off. The complete example recipe is the runnable version.
What this gives you is a fast restart. It is not a standby: SGLang has no election, so the blog's "active-standby" scenario is a deployment pattern, not a feature. For takeover by a parked standby engine, see shadow engine recovery for vLLM through Dynamo's GPU Memory Service.
Table of Contents¶
- Quick Start
- What Runs
- What srtctl Owns vs What You Set
- Testing a Relaunch
- Limitations
- Troubleshooting
Quick Start¶
This is an excerpt of the weight-cache settings, not a complete recipe. Start from the complete example recipe, which includes the required name, model path, precision, and resources, and resolve its model and container aliases in your cluster config. Each worker in this recipe fits on one node.
schema: 2
model:
container: lmsysorg/sglang:v0.5.20 # v0.5.19 or newer
frontend:
type: sglang-router
engine: sglang
roles:
agg:
nodes: 1
workers: 2
gpus: 2
restart: {policy: always, backoff_seconds: 5}
env:
SGLANG_WEIGHT_CACHE_SOCKET_TEMPLATE: "/dev/shm/sglang-wc/{device_uuid}.sock"
SGLANG_WEIGHT_CACHE_READY_TEMPLATE: "/dev/shm/sglang-wc/{device_uuid}.ready"
args:
tensor-parallel-size: 2
weight-cache-mode: client
services:
- name: weight-cache
type: generic
start: before_workers
critical: true
placement: {node: agg, per: worker}
env:
SGLANG_WEIGHT_CACHE_SOCKET_TEMPLATE: "/dev/shm/sglang-wc/{device_uuid}.sock"
SGLANG_WEIGHT_CACHE_READY_TEMPLATE: "/dev/shm/sglang-wc/{device_uuid}.ready"
preamble: "mkdir -p /dev/shm/sglang-wc"
command: ["python3", "-m", "sglang.srt.weight_cache.daemon", "--model-path", "/model", "--tensor-parallel-size", "2"]
readiness:
log: {pattern: "All 2 weight cache daemons on node 0 are ready"}
timeout_seconds: 900
srtctl dry-run -f <recipe> lists the service with its placement, probe and environment before you submit.
What Runs¶
For each single-node TP2 worker in the example:
sequenceDiagram
participant S as srtctl supervisor
participant C as Per-worker weight-cache service
participant E as SGLang engine
S->>C: Launch in the worker's CUDA_VISIBLE_DEVICES
C->>C: Load weights into HBM, one daemon per TP rank
C-->>S: All 2 weight cache daemons on node 0 are ready
S->>E: Launch with --weight-cache-mode client
E->>C: Connect to GPU UUID sockets under /dev/shm/sglang-wc
C-->>E: Validate fingerprint and export CUDA IPC handles
E-->>S: Engine exits
S->>E: Relaunch after backoff, same node, GPUs and ports
E->>C: Reconnect to the same GPU UUID sockets
C-->>E: Map the same tensors without reading the checkpoint
- The service.
placement.per: workergives one instance per worker on each of its nodes, pinned to that worker's device mask, so the daemon's device k and the engine's device k are the same GPU and derive the same UUID.start: before_workersputs it up before the engines and, at cleanup, stops it after them (shutdown tier 1). It iscriticalbecause an engine's liveness watchdog SIGKILLs the engine when its daemon PID disappears (the mapped pointers would dangle). - The paths. Both sides read
SGLANG_WEIGHT_CACHE_SOCKET_TEMPLATEandSGLANG_WEIGHT_CACHE_READY_TEMPLATE; the template must contain{device_uuid}, which srtctl's placeholder rendering leaves alone. SGLang's default is/tmp, a per-container tmpfs under enroot that the engine step cannot see./dev/shmis the host's, bind-mounted into every container on the node. One node-wide directory is enough because the files are keyed by physical GPU UUID; the daemon removes its files on SIGTERM and cleans up stale ones (dead PID) when it starts. - The engine.
--weight-cache-mode clientconnects to the daemon for its GPU, validates a fingerprint (model path and architecture, TP/PP/DP/EP ranks, quantization method and config hash, dtype, compute capability, torch version), and maps every tensor zero-copy. An absent socket falls back to a disk load; a refused connection or a fingerprint mismatch raises. - The relaunch. With
restart: alwaysoron-failure, the supervisor relaunches an exited engine step in place (same node, GPUs and ports,_r<n>step name, crash log rotated aside) after the backoff. The relaunched engine maps the daemon's tensors again.worker_restarts.jsonand the lockfile record each event.
What srtctl Owns vs What You Set¶
| Piece | Owner | Value |
|---|---|---|
| one daemon group per worker, in its device mask | srtctl | placement: {node: <role>, per: worker} |
| start order and shutdown order | srtctl | start: before_workers (tier 1, stopped after the engines) |
| readiness gate | you | readiness.log.pattern with the rank count the daemon prints (All <tp> weight cache daemons on node <rank> are ready) |
| the daemon's argv | you | the engine's model, parallelism, dtype and quantization flags, exactly; a mismatch is refused at engine start, not at load |
| socket and ready paths | you | the two templates, identical on the service and in roles.<role>.env, under /dev/shm |
| relaunch | you | roles.<role>.restart; without it an exited engine's step just ends and critical decides |
| the image | you | SGLang v0.5.19 or newer (env-configurable paths, GPU-UUID keyed sockets, static DP/EP layouts) |
Testing a Relaunch¶
Kill the engine process, not its step. Use your job ID and the worker's node and allocated HTTP port from the sweep log's Command: line. From the login node:
srun --jobid <job> --overlap -w <node> -N1 bash -c 'pkill -9 -f "sglang.launch_server.*--port <worker-port>( |$)"'
Then watch:
- the sweep log:
Worker agg_0_<node> exited with code 137; relaunching agg_0 in 5s (restart 1/3), thenRelaunched agg_0 as agg_0_<node>_r1; - the new engine log (
<node>_agg_w0.out; the crash log moved to<node>_agg_w0.out.1):[IpcModelLoader] Loaded model via IPC (mode=client)andLoad weight end. elapsed=0.0x s, thenThe server is fired up and ready to roll!; - the router's
/workerson the head node: the worker's URL is unchanged, so it is back in rotation as soon as it answers.
The daemon service's log shows nothing during a relaunch; the tensors stay mapped in the daemon regardless of how many engines come and go.
Limitations¶
- Fast restart only. No standby engine and no cutover; see shadow engine recovery for that on vLLM.
- Restarts have a finite per-worker budget. Once it is exhausted, the role's
criticalsetting decides whether the run fails; see restart behavior and the generated RestartPolicy reference for configuration fields and defaults. - SGLang v0.5.19 or newer. Older builds hard-code
/tmp/sglang_weight_cache_rank*paths, which a separate srun step cannot reach. - IPC-safe quantizations only: unquantized and block-wise FP8 as of v0.5.20. Per-tensor FP8, Marlin, AWQ/GPTQ and NVFP4 raise at daemon start.
- Not with speculative decoding (
--speculative-algorithm); the daemon does not export the draft model. - The daemon's argv is written by hand. It must carry the same model, parallelism, dtype and quantization flags as the engine; srtctl does not derive one from the other.
- Workers spanning multiple nodes are not supported by this readiness-gated recipe. srtctl launches each per-worker service instance and waits for its readiness before launching the next. The daemon's distributed rendezvous needs all nodes to start before any node reports ready, so the first node times out while the other nodes are still unlaunched. Adding
--nnodes,--node-rank, and--dist-init-methodalone does not fix this. Multiple independent single-node workers are supported. - Per-worker service instances start one after another, each gated on its readiness probe, so a node with N workers pays N daemon start-ups in sequence.
- The daemon is not relaunched. If it dies, its engines SIGKILL themselves (dangling mappings) and, with a restart policy, come back through a disk load or fail on a refused socket;
critical: trueon the service fails the run instead.
Troubleshooting¶
| Symptom | Cause | Fix |
|---|---|---|
engine log: Daemon socket not found at /tmp/sglang_weight_cache_... and a normal disk load |
the templates are not set on the engine, or point at a per-container path | set both templates in roles.<role>.env under /dev/shm |
service log: sglang serve: error: unrecognized arguments |
the daemon parses the full engine argument set; a flag the image no longer knows | drop the flag; the same flag would break the engine too |
engine log: Daemon config mismatch |
the daemon's argv differs from the engine's (model, TP, dtype, quantization) | mirror the engine's flags in services[].command |
engine dies with Weight cache daemon (pid=...) died while this engine holds its weights |
the service was stopped or crashed | keep start: before_workers (stopped after the engines) and critical: true; check the service log |
readiness probe times out while the daemons say Listening on ... |
the pattern's rank or node count does not match the launcher's summary line | use the printed All <n> weight cache daemons on node <k> are ready text |
already running (pid=...) at daemon start |
a live daemon already owns that GPU's socket | confirm which job owns the daemon before stopping it; remove a leaked daemon only after verifying its job has ended |