Skip to content

Observability and Telemetry

The observability: preset (nsys, OTEL, Tachometer) and the telemetry: providers (GPU and CPU power). Profiler settings: Profiling.

observability

Tachometer collection is on by default for every run (no configuration needed; observability.tachometer.enabled: false opts out). observability.enabled turns on the server metrics content (the TRT-LLM publish flag and engine statistics) and the trace surfaces:

observability:
  enabled: true

The capture window aligns with the load, the same window the benchmark client's own AIPERF_SERVER_METRICS_URLS polling covers: on benchmark runs the scraper starts once the server passes the health gate (bring-up produces only dead-endpoint noise while workers load) and is stopped gracefully when the client exits, with a configurable grace period for compacting final.parquet before post-processing reads it. Runs without a discrete load window (serve-only, manual, eval-only) capture the whole serve session as before. Signal handlers and the critical-process monitor still use the process registry's existing teardown budget; the benchmark shutdown grace does not override those paths.

Tachometer scrapes all configured worker, frontend, DCGM, and node-exporter endpoints, independently of the benchmark client's AIPERF_SERVER_METRICS_URLS polling. This keeps the raw capture complete even when the client also collects metrics.

The legacy in-job Python RAW scraper is retired: a recipe still carrying scrape_metrics, scrape_interval_seconds, or scrape_output fails validation at submit time. Historical raw_prometheus.jsonl artifacts remain readable by the post-processing ingest.

Fields: ObservabilityConfig, NsysObservabilityConfig. tachometer.enabled: null (the default) collects on every run, independent of observability.enabled; explicit false opts out. See Profiling for the nsys settings.

observability.enabled: true also enables nsys with NVTX tracing and CPU sampling on every frontend and worker process/rank. Its default nsys.capture_window: measured_workload starts collection after warmup and stops it when the measured workload finishes. Supported bundled runners call the boundary hooks automatically; custom/manual clients must call them at their own phase boundaries. Set nsys.capture_window: including_startup to include startup and warmup through teardown, or nsys.enabled: false to opt out. An enabled top-level profiling mode takes precedence. The serving container must provide nsys and the required NVTX support. See Observability capture for timing, sampling, injection, and report-finalization settings.

The component perf dashboard is built explicitly after a run; enabled decides which capture legs exist and therefore which tabs a later build carries. Jobs do not automatically run dashboard ingestion or rendering. See Component Performance Dashboard.

SGLang workers always receive --enable-metrics unless the recipe sets it: native sglang.launch_server serves /metrics only with the flag, and dynamo.sglang merges the engine's sglang:* series into its system-port /metrics only when SGLang was started with it.

Tachometer collects every worker rank, frontend, DCGM, node, and process metrics by default (minus the client-polled complement described above); the exporters launch from pinned multi-arch registry images with no configuration. Air-gapped clusters override the images via the containers: alias map in srtslurm.yaml; default_exporters: false disables the built-ins:

observability:
  enabled: true
  tachometer:
    enabled: true
    collect_interval_ms: 1000
    sync_interval_secs: 120
    compaction_threads: 4
    storage_subdir: tachometer
    extra_metadata:
      cluster: production
    dcgm_exporter:
      container_image: /containers/dcgm-exporter.sqsh
      port: 9400
    node_exporter:
      container_image: /containers/node-exporter.sqsh
      port: 9100
    process_exporter:
      binary: /opt/srt/configs/process-exporter   # host-native (default mode); or set container_image instead
      container_image: ""
      port: 9256

Every Tachometer field: TachometerConfig. The ones whose behavior needs more than a line:

Tachometer field Type Default Description
collect_interval_ms int 1000 Milliseconds between scrapes of every endpoint; the single cadence knob. It also drives the launched DCGM exporter's --collect-interval (an explicit dcgm_exporter.command wins) and the host sampler. Values below 1000 speed up DCGM NVML sampling and are warned about at launch: 100ms sampling measured ~2% decode ITL overhead on GB300. Replaces the retired Hz-based default_frequency
shutdown_grace_secs float 120.0 Time the scraper gets after SIGTERM to flush and compact final.parquet before it is killed; compaction scales with the data accumulated since the last periodic sync
default_exporters bool true Imply the built-in DCGM + node + process exporters when no explicit blocks are set. The exporters are services (dcgm-exporter, node-exporter, process-exporter); declaring one under services: by that name overrides it, and srtctl dry-run lists them
dcgm_exporter object/null built-in Defaults to nvcr.io#nvidia/k8s/dcgm-exporter:3.3.9-3.6.1-ubuntu22.04 on port 9401; an explicit block overrides
node_exporter object/null built-in Defaults to quay.io#prometheus/node-exporter:v1.8.2 on port 9101 with the cpu, infiniband, meminfo, processes, stat, vmstat, pressure and meminfo_numa collectors on worker nodes. Includes major faults and page-reclaim counters; retains process-state and NUMA-node identity. An explicit block overrides
process_exporter object/null built-in Defaults to the host-native configs/process-exporter binary (ncabatoff/process-exporter 0.8.7, installed by make setup for the compute arch, like configs/nats-server and configs/etcd) on port 9256, launched with plain srun (no container) on every allocated node (the process-exporter service, placement.node: all). Reads the host /proc and publishes per-process-group CPU seconds by mode, thread count, per-thread-name CPU and count (-threads=true), context switches, RSS and open fds. The passthrough filter retains metric names and labels, including summary sum/count suffixes, and attaches hostname and run metadata to raw rows. Groups (frontend, dynamo_trtllm / dynamo_sglang / dynamo_vllm handlers, trtllm_engine children, the SGLang engine processes sglang_scheduler / sglang_dp_controller / sglang_detokenizer matched on their retitled command lines, launcher, client, infra daemons) come from <log_dir>/process-exporter.yml, written at launch. If the binary is missing the leg is skipped with a warning (submit warns too). An explicit block may set binary (absolute, or relative to the srtctl checkout) or instead a container_image with binary unset to run it containerized; the upstream FROM scratch image is not used by default because pyxis/enroot on some clusters cannot start shell-less images

Every exporter block (TelemetryExporterConfig) accepts container_image, port, command and binary. binary selects host-native launch (the executable runs directly under srun, container_image is ignored and may be ""); without it the exporter runs from container_image. One of the two must be set.

make setup ARCH=<compute_arch> downloads and checksum-verifies the matching Tachometer binary from the latest srt-slurm release and installs the process-exporter binary for the same arch. The scraper and the process exporter run as native srun processes; the DCGM and node exporters remain containerized on worker nodes. Run make tachometer-scraper to build the scraper from source instead. The process-exporter passthrough filter and node process-state/NUMA-label preservation require a scraper built from this revision or a release containing it; rebuild the scraper when using an older downloaded binary.

The pressure collector reports PSI only when the host exposes the corresponding /proc/pressure files; missing metrics indicate unavailable data. NUMA memory and allocation metrics retain the exported node label as numa_node in raw metric names, separately from host metadata. With observability.enabled: true, the existing local host sampler also records cumulative PSI stall totals in microseconds in its psi JSONL field. That optional sampler covers the sweep/orchestrator host only; it does not extend exporter placement to dedicated frontend or client nodes. Collector overhead has not been measured for this change.

Tachometer writes its Parquet stream under <log_dir>/<storage_subdir>/raw/scrape/ (the leaf is created by the scraper itself; srtctl pre-creates only the parent, because the scraper refuses a pre-existing storage directory), compacting to final.parquet there on shutdown. Intermediate files remain in <log_dir>/<storage_subdir>/local until shutdown compaction completes. Rows carry an epoch timestamp_ns column, so they join directly with AIPerf records and Dynamo spans; explicit dashboard ingestion converts the Parquet into the dashboard's server_metrics_export.jsonl.

The scraper runs as a best-effort process: if it dies (or the binary is missing at runtime), the benchmark continues and the loss is visible in tachometer.out and the sweep log. srtctl validate-setup still fails fast at submit time when bin/tachometer-scraper is absent.

telemetry

telemetry is reserved for DCGM power measurement. It can run alongside observability.tachometer; it does not start Tachometer itself.

When both are enabled, telemetry.dcgm_exporter is shared with Tachometer. Do not also configure observability.tachometer.dcgm_exporter; Tachometer can still launch an optional node exporter from its own block.

telemetry:
  enabled: true
  collect_interval_ms: 1000
  storage_subdir: power
  required: true
  dcgm_exporter:
    container_image: /containers/dcgm-exporter.sqsh
    port: 9400
  cpu_power_exporter:
    port: 9405
    source: auto

Fields: TelemetryConfig. collect_interval_ms is shared by the DCGM and CPU legs and must be at most 3000.

telemetry requires a benchmark.type of sa-bench, custom, agentic, agentx, or manual (a manual job has no load window, so like serve-only it captures the whole serve session; use it when an external load generator drives the endpoint), the benchmark client on the head node (benchmark.placement.node: head, the default), and no dedicated node for the discovery plane (an etcd/nats service with placement.node: dedicated moves the head off the batch host the collector runs on).

CPU power

telemetry.cpu_power_exporter is an independent, best-effort leg: its presence (not a separate enabled flag) turns CPU power collection on, and it can run with or without dcgm_exporter alongside it. On each worker node, srtctl launches a cpu-power-exporter process directly on the bare host (outside the model container, so it can read host power interfaces) and exposes it on cpu_power_exporter.port. It resolves the bundled Rust binary installed by make setup first, falling back to the ACPI-only Python stdlib exporter (srtctl.core.cpu_power_exporter) when that binary is absent. A head-node collector scrapes every worker's exporter on the shared collect_interval_ms/request_timeout_seconds cadence and writes per-sample rows plus a manifest under <log_dir>/<storage_subdir>/cpu/ (samples.csv, cpu_manifest.json).

Fields: CpuPowerExporterConfig. source: auto tries DCGM first and falls back to ACPI; it has no effect on the Python fallback exporter, which is ACPI-only.

samples.csv carries one row per sensor reading, with columns schema_version, timestamp_unix, hostname, source, sensor, socket_id, power_w, total_power_w. In ACPI mode, total_power_w is not a sum of the cpu- and sysio-kind rails; whenever a grace-kind channel exists for a socket, that channel alone is the node-level total (real hardware traces show grace at roughly 93-104W against cpu+sysio combined at roughly 53-58W for the same socket, i.e. grace measures the whole Grace SoC power boundary, not literally cpu + sysio). When no grace channel is present for a scrape, total_power_w is left blank for that row rather than guessed from the component rails; per-socket power_w values are always populated regardless. In DCGM mode, total_power_w is the single already-aggregate value DCGM reports per socket.

Because collection is always best-effort, there is no required knob for the CPU leg: a node that fails to expose its exporter (or fails to publish readings) simply produces gaps in samples.csv, and the run continues.