Skip to content

GPU Power Telemetry (dcgm-power)

The dcgm-power telemetry provider records raw per-GPU watts for every allocated worker node, the topology needed to map each GPU to a prefill, decode, or agg role, and the exact formal benchmark window for every measured concurrency. It never integrates power into energy and never branches on model, precision, or recipe; consumers integrate watts over the recorded window themselves.

How it works

  • One DCGM exporter task runs on each allocated worker node, launched through the normal SLURM/process-registry path (one srun per heterogeneous group).
  • A collector thread inside the orchestrator polls every exporter concurrently from the physical head node, so all sample timestamps and benchmark boundaries come from one clock.
  • Only DCGM_FI_DEV_POWER_USAGE is parsed. Device identity comes from the gpu and UUID labels.
  • No in-tree benchmark stamps measurement windows yet, so every run is currently unpublishable: it records MEASUREMENT_WINDOW reason codes, and required: true exits non-zero. The adapter belongs with the benchmark (for the current sa-bench path and its planned replacement alike): the benchmark child writes one window file per measured concurrency using the standalone measurement_window.py module and the windows directory passed in via MEASUREMENT_WINDOW_DIR_ENV.

Configuration

# NOTE: unsupported end-to-end until a benchmark adapter stamps windows —
# with this exact config every run is unpublishable and `required: true` fails.
benchmark:
  type: sa-bench          # future benchmark-side adapter must stamp the windows
  placement:
    node: head            # keeps sample and window clocks on one host
  isl: 8192
  osl: 1024
  concurrencies: [4]

telemetry:
  enabled: true
  collect_interval_ms: 1000         # milliseconds between collector cycles; must be <= 3000
  storage_subdir: power             # relative to the run log directory
  required: true                    # exit non-zero when artifacts are unpublishable
  startup_timeout_seconds: 30
  request_timeout_seconds: 2
  collector_join_timeout_seconds: 12
  dcgm_exporter:
    container_image: dcgm-exporter  # alias, path, or registry URI
    port: 9401

dcgm-power needs only dcgm_exporter: there is no provider key, and it does not require the top-level container_image or a node_exporter, because the collector runs inside srtctl. Config loading validates the block and rejects inconsistent values with actionable messages; in particular collect_interval_ms must not exceed the 3-second max sample gap the validator accepts, or every window would fail sample_gap_exceeded. Telemetry stays disabled by default. The collector join timeout must exceed two complete request-cycle budgets (2 * (2 * request_timeout_seconds + 1 second)), covering a scrape already in flight when shutdown starts plus the final bracketing scrape.

Artifacts

<log_dir>/<storage_subdir>/
├── manifest.json
├── samples.csv
└── windows/
    └── <benchmark-result-stem>.json

samples.csv has the exact header schema_version,timestamp_unix,scrape_seq,hostname,gpu_index,gpu_uuid,power_w, one row per observation, (scrape_seq, hostname, gpu_index) unique. Rows are never interpolated, averaged, or role-attributed — role and heterogeneous group live once in the manifest topology.

manifest.json records producer identity (version, git commit, exporter image and its SHA-256), the sample interval, expected and observed device sets, the topology mapping, the expected window list, the SHA-256 of the finalized samples.csv bytes, terminal status, per-window coverage validation, and reason codes. status is the lifecycle outcome; publication_valid is the separate publication gate. Reason codes are stable machine-readable strings enumerated in srtctl/core/power/contract.py. The digest is required for offline publication validation, so packages created before samples_sha256 was recorded cannot be certified by this validator.

A window file records the formal benchmark boundaries on the head-node Unix clock plus a monotonic duration, and points at the SA-Bench result it brackets; result and window are boundary-identical.

With required: true, all artifacts are written first and the job then exits non-zero whenever the terminal manifest is not publishable. With required: false, measurement invalidity leaves the benchmark exit code unchanged; an operational failure — a collector that cannot be joined, a benchmark child that cannot be reaped, or an internal error while finalizing telemetry — fails the job in either mode.

On SIGTERM/SIGINT or a critical-process death, the shared process registry tears processes down before the collector finalizes, so the final scrape sees dead endpoints. The manifest fails closed (exporter_exited / collector_interrupted force publication_valid=false); the cost is that a job that was simply cancelled can record exporter_exited.

Re-validating a retained run

The artifact package is self-describing. The manifest supplies producer identity, expected topology, runtime-only failure history, and a stored verdict; the validator does not trust that verdict on its own:

srtctl-validate-power \
  --power-dir outputs/12345/logs/power \
  --result-root outputs/12345/logs \
  --expect-role prefill=4 --expect-role decode=4 \
  --require-distinct-het-groups

It recomputes every disk-derived claim from samples.csv, the result files, and windows/, then requires the stored disk-derived reason subset and publication_valid verdict to agree. Runtime-only reasons such as HTTP, exporter-process, and collector failures cannot be reconstructed after the live job is gone, so they are checked for a known v1 enum value and lifecycle consistency instead. Exit status is 0 only when the recomputed package is publishable, the stored verdict is true, and the two agree; otherwise it is 1 and every failure is printed. The --expect-* flags optionally assert an expected job shape for hardware canaries.

Cumulative sample coverage

Each expected GPU must also retain at least 95% of the expected sampling intervals across its nearest bracketing samples. Expected intervals are floor((last_bracket - first_bracket) / sample_interval_seconds); observed intervals are the number of distinct sample times in that same span minus one, so a row written twice counts once. This avoids counting ordinary cadence jitter as repeated loss or allowing warmup samples to hide missing measurements. More than 5% missing intervals, more than one sample in twenty, records sample_loss_exceeded, even when every individual gap is below 3 seconds. For example, sampling every 2 seconds with a recorded 1-second cadence fails. A manifest without a finite positive sample_interval_seconds fails this rule too, alongside the manifest field check.

The 3-second maximum gap and boundary checks still apply, but on short spans the loss rule is the stricter one. At a 1-second cadence one dropped sample passes only from 20 intervals, and one 3-second hole (two missing intervals) only from 40; a shorter window rejects on that single loss by design, because one lost second is a larger share of it. Formal sa-bench windows run for minutes, where a single gap at the limit passes both rules. Session finalization and offline validation use the recorded cadence and the same coverage rule. Previously accepted sparse packages can fail revalidation; their files are not rewritten. This limits sample loss, not the numerical error in energy.

Diagnosing slow scrapes

The collector writes best-effort scrape-timings.jsonl beside samples.csv. Every line carries an event: scrape per settled endpoint request, cycle_write per collection cycle, and one closing diagnostic_summary. Join a scrape record to its GPU rows using (hostname, scrape_seq). Each records its start/end times, HTTP status or exception, request and parse durations, sample timestamp, row count and reason codes. Failed HTTP requests retain timing records, with null parse duration and sample timestamp, without inventing power samples. Requests still unsettled when the cycle deadline expires have no timing record.

Instants are unix timestamps; durations come from the monotonic clock. schedule_lag_seconds measures request start against the background cycle's scheduled slot, which cycle_write records as scheduled_at_unix; manual and final bracketing scrapes use null for both. The collector writes a cycle's endpoints together, so the cycle_write record keyed by scrape_seq carries the batch's writer_lock_wait_seconds, sample_write_seconds and attempted row_count once. Its sample_write_completed reports whether the batch was appended and flushed; when it is false, sample_write_error names the exception class if the append raised, or is null when the session was already finalizing and refused the batch.

Only a daemon writer performs diagnostic file I/O, outside the sample writer lock. Its queue holds at most 128 pending records; overflow drops diagnostics, not power samples. A final diagnostic_summary reports dropped_records. Missing summary means diagnostics may be incomplete. Shutdown waits only until the existing collector deadline. This optional sidecar is not publication validation evidence, and its absence or write failure does not invalidate power.