Skip to content

DSight metric sources and log generators

The Metrics panel reads one normalized series/catalog schema. Its inputs are Tachometer Parquet/Arrow captures and supported worker logs. Both go through the same series finalization, source evidence, compressed family loading, charts, pinning and query APIs. No AIPerf metric summary is used for these families.

Source to UI

flowchart LR
    T["Tachometer Parquet / Arrow rows"] --> TR["metrics.read_metrics<br/>names, labels, absolute timestamps"]
    W["Worker .out lines + filename identity"] --> G["LogMetricGenerator base class<br/>TokenSpeed and SGLang adapters"]
    G --> LR["log_metrics.reader<br/>timezone, scope, file + line evidence"]
    TR --> S["Shared metric series + catalog<br/>deduplication, conflicts, reference IDs"]
    LR --> S
    S --> D["trace-data.json.gz<br/>Python / CLI / MCP queries"]
    S --> H["index.html<br/>lazy compressed JSON families"]
    H --> UI["Metrics panel<br/>pinned charts, legend, sampled peaks<br/>solid observations + dashed limits"]

The browser does not parse Parquet. Log-derived metrics enter the shared normalized schema directly; no intermediate Parquet conversion is required.

Tachometer capture schema

Field Type Meaning
timestamp_ns int64 Absolute UTC epoch nanoseconds; required for alignment
metric_name string Metric name with optional inline labels
metric_value float64 Captured scalar or histogram bucket value
scraper_endpoint string Source endpoint identity
time_since_start float64 Tachometer-relative time; DSight aligns on timestamp_ns
histogram_bucket_lower, histogram_bucket_upper, histogram_sum, histogram_count nullable float64 Histogram capture fields
Additional metadata columns strings Host, worker, GPU and other source identity dimensions
metric_name_clean string, when present Compaction-derived name; not required by DSight

The first four columns are required by the DSight metric reader. Histogram bounds contribute to series identity; sum/count columns do not create additional series. final.parquet supersedes compacted shards, while an Arrow tail can still be read. The reader detects Parquet versus Arrow from contents, including captures whose file suffix does not match their encoding.

Normalized series

Each series carries id, name, label, unit, catalog presentation metadata, worker, host, gpu, rank, rank_kind, worker_process, labels, source_ids, and points. A point is:

[elapsed_seconds_from_run_origin, value, source_id, source_row_or_line]

source_id indexes the dataset's source manifest. Tachometer rows are zero-based; worker log lines are one-based. The manifest retains the file identity and hash. Log series additionally declare source_kind: "worker_log", generator, time_resolution_s, and temporal:

  • sample: a recorded observation; no value is assumed before or after it.
  • event: a recorded per-request observation; distinct source lines remain distinct points even when timestamp and value match. Simultaneous events do not count as conflicting metric samples.
  • setting: a recorded configuration held until the next setting in the same scope. A null value invalidates the previous setting. The last timestamp group before the trace window is retained at its original, possibly negative, elapsed timestamp. Queries expose preceding-window evidence as carried_setting, separately from in-range points and sample statistics.

An observed series can declare a reference with name, label and series_id. The reference is another ordinary metric series in the same unit, independently selectable and queryable. A missing match has series_id: null.

Reference matching requires the same source file, worker, rank namespace/rank, recorded process and labels. It never pools ranks, borrows another worker's configuration, or equates attention TP rank with GPU/global/DP rank. When a log has no process identity, that field remains null; file and rank still isolate it. Conflicting values at one timestamp remain in raw evidence and appear as gaps.

Dynamo–TokenSpeed metrics

Metric name to select or pin Source field Unit Paired reference
log_tokenspeed_active_decode_requests Decode batch #running-req requests log_tokenspeed_decode_request_limit
log_tokenspeed_decode_request_limit Scheduler config max_batch_size requests —
log_tokenspeed_active_kv_pages #pages(active/cached/total) → active pages log_tokenspeed_kv_pool_pages
log_tokenspeed_kv_pool_pages #pages(active/cached/total) → total pages —

These families appear under Workers / Log-derived metrics. Generated metrics use log_<component>_<name>, where <component> identifies the component that produced the consumed log. The log_ prefix distinguishes generated evidence from native Prometheus metrics. These logs come from TokenSpeed, so their component is tokenspeed, even when TokenSpeed runs through the Dynamo integration. Active decode batch is not the exported tokenspeed:num_requests_running scheduler-state count. The configured per-scheduler batch limit is not global max_num_seqs or benchmark concurrency. KV pool size comes from the same snapshot as active pages, not num_device_pages, reserved-page arithmetic, or token counts.

The adapter reuses the existing TokenSpeed batch decoder and recognizes scheduler configuration lines. It preserves attention TP rank. Local timestamps need an explicit --iteration-timezone; otherwise these metrics are omitted with a warning. Missing page fields or configuration do not create zero-valued samples. A new scheduler configuration without a valid positive max_batch_size clears the previous limit. Settings are not applied before their recorded timestamps.

SGLang batch metrics

SGLang Prefill batch and Decode batch lines, with or without a [counter], contribute the recorded fields below as log_sglang_* observations. The shared reader retains each physical log line as evidence. Each series keeps the DP rank when logged, plus any pp, attn_cp, moe_dp, tp, and ep labels and a phase label (prefill or decode); values from different ranks, phases, files or workers are never summed. A field absent or invalid on a line creates no point. All values are snapshots or logger-reported rates, not held settings or continuously measured occupancy.

Phase Recorded fields Metric suffixes Units
Prefill #new-seq, #new-token, #cached-token new_sequences, new_tokens, cached_tokens requests, tokens, tokens
Prefill #pending-token, #bootstrap-req, #inflight-req, #optimistic-req pending_tokens, bootstrap_requests, inflight_requests, optimistic_requests tokens, requests, requests, requests
Prefill input throughput (token/s) input_throughput_tokens_per_second tokens/s
Decode #token, #prealloc-req, #transfer-req, #retracted-req decode_tokens, preallocated_requests, transfer_requests, retracted_requests tokens, requests, requests, requests
Decode accept len, accept rate, pre-allocated usage, gen throughput (token/s) accept_length, accept_rate, preallocated_usage, generation_throughput_tokens_per_second tokens, ratio, ratio, tokens/s
Both #running-req, #queue-req, token usage, cuda graph running_requests, queued_requests, token_usage, cuda_graph_enabled requests, requests, ratio, boolean (0/1)

The suffixes in this table have the log_sglang_ prefix in the catalog. #token is the decode pool's used-token count excluding available and evictable tokens. SGLang logs #cached-token as batch log_hit_tokens; it is not a full-workload cache hit rate. token usage is the reported token-pool ratio, and pre-allocated usage is the preallocated-token count divided by the scheduler token capacity. accept len and accept rate are the speculative acceptance values reported for that decode logging interval. Input and generation throughput are the logger's own rates over its preceding logging interval. The adapter does not derive a cache-hit fraction, resample a rate, or infer values between samples.

SGLang timestamps are local. Pass the run's actual timezone through --iteration-timezone to align them with client and exported telemetry; without it the reader omits these metrics and records a warning. Both default second-resolution timestamps and optional fractional seconds are supported. SGLANG_LOG_MS and SGLANG_LOG_FORWARD_ITERS are not required. Rank prefixes may contain any of DP, PP, ATTN_CP, MOE_DP, TP, and EP, or no ranks. Without a logged DP, the normalized rank and rank kind remain unknown; a TP label is not reinterpreted as a DP rank. A series records the coarsest timestamp precision among its observations (1 s without a fraction). These formats follow the pinned upstream logger, rank prefix, and batch formatter.

Batch observations use event semantics because multiple batches can share one logged second. Queries retain every line, including repeated values; the chart shows a labeled median and event count for coincident observations. It does not fabricate ordering or subsecond timestamps. The synthetic SGLang example includes both default and detailed formats and builds a report from these logs alone.

Per-request timing records

SGLang ReqTimeStats(...) lines also enter the same catalog. The request's rid stays in the source line, not in series labels. A request emits one record when the scheduler sees it finished and request-time logging is enabled; the point is placed at the log emission time, which is an observation after completion, not the queue-entry or first-forward time. Scope retains the logged rank labels when present and type=prefill or type=decode as the phase label.

Metric suffix after log_sglang_ Source or calculation Unit
request_input_tokens, request_cached_input_tokens input_len, cached_input_len on that request tokens
request_uncached_input_tokens input_len - cached_input_len, when both counts are valid and cached ≤ input tokens
request_cached_input_fraction cached_input_len / input_len, when input > 0 and 0 ≤ cached ≤ input ratio
request_bootstrap_duration_ms, request_queue_duration_ms, request_forward_duration_ms Corresponding ReqTimeStats durations on either phase ms
request_allocation_wait_duration_ms, request_transfer_duration_ms Decode-only alloc_wait_duration, transfer_duration ms
request_bootstrap_queue_duration_ms Prefill bootstrap_queue_duration when bootstrap has not completed ms
request_preallocation_queue_duration_ms Decode prealloc_queue_duration when no bootstrap-completion timestamp is available ms
request_transfer_speed_gib_per_second, request_transfer_total_mib Prefill-only transfer fields as logged GiB/s, MiB

These durations describe the individual request's recorded stages. They cannot be summed across requests, equated with client latency, or assigned to the emission timestamp as if that were their start. request_cached_input_fraction is a per-request fraction, not the batch's #cached-token share or a full-workload cache hit rate. A missing or invalid input count leaves that derived value absent; a valid zero denominator remains unknown rather than zero. Request metrics use event semantics: queries and SQLite keep every request's point and source line, including repeated values at the same logged timestamp. The chart displays the median when events share a timestamp and labels the number of events at that tick. This display value is not an additional raw sample; hover and the source query distinguish it from individual requests. The line between event ticks is a visual guide, not continuous occupancy. SGLang computes transfer_total by dividing bytes by 1024² and transfer_speed by 1024³, despite printing MB and GB/s; the catalog uses their binary units. For chunked prefill transfer, the logged transfer timing can cover only the last chunk, so these fields do not establish a whole-request transfer rate. Its duration formatter returns zero when either timing endpoint is unavailable, so a logged zero does not always prove a zero-length stage. Prefill forward_duration runs from forward entry through completion and can include chunking and transfer. The logged prefill entry_time is the bootstrap queue entry time, while queue_duration starts at the waiting queue entry; the adapter does not use entry_time as an alignment anchor. Alternative queue durations retain their own metric names; they do not create request_bootstrap_duration_ms or request_allocation_wait_duration_ms points. When prefill instead logs bootstrap_done_time, that wall-clock timestamp stays in the source evidence and is not converted into a duration. Other valid fields on the same request still contribute points. The upstream request-timing implementation defines the stage boundaries, missing-endpoint behavior, and binary transfer units.

Capacity presentation

Pin Active decode batch and Active KV pages to see activity and capacity on the same time axis. Each source has a solid observed line and a matching dashed limit in the same units. Hiding that source hides both lines. Capacity charts start at zero and include the reference in the vertical scale.

The legend highlights Peak observed and the limit/pool value using exact counts. Highest observed usage is the maximum of each sample divided by its corresponding valid positive limit. It is not the ratio of unrelated maxima, a time-weighted average, or proof of continuous saturation. Zooming recomputes these summaries for the selected time range, separately for each source.

A configuration reference becomes a dashed step, including when its establishing log precedes the selected window. Only display coordinates are extended to view boundaries; source samples and their timestamps are unchanged. A sampled pool reference is bounded by its samples; usage requires a matching sample timestamp. Changed limits display a range marked changed. Missing/ambiguous limits show unavailable, including the count of observed samples without a valid limit. The observed metric still works without a reference, Tachometer, OTel or Nsight.

Generator interface

log_metrics/base.py defines frozen LogMetricDefinition and LogMetricEvent records and the LogMetricGenerator abstract base class:

from abc import ABC, abstractmethod

class LogMetricGenerator(ABC):
    @property
    @abstractmethod
    def name(self) -> str: ...

    @property
    @abstractmethod
    def definitions(self) -> tuple[LogMetricDefinition, ...]: ...

    @abstractmethod
    def parse_line(self, line: str, source: SourceIdentity) -> LogMetricEvent | None: ...

A generator returns timestamped values and recorded rank/process/label scope; unsupported lines return None. Definitions supply metric names, units, sample versus event/setting semantics and optional reference relationships. The shared reader owns file discovery, timezone alignment, window selection, evidence, validation, normalization and reference matching. Configuration-only evidence does not invent a workload time envelope for a source-only report.

log_metrics/__init__.py holds the generator registry. Implementations inherit LogMetricGenerator and provide all three abstract members; class attributes can supply name and definitions. Add each implementation to the registry with representative source fixtures and missing/changed-limit tests. Engine log syntax stays in the adapter; rendering depends only on the normalized contract. DynamoTokenSpeedLogMetrics and SGLangLogMetrics are registered generators.

Focused checks:

uv run pytest tests/test_dsight_log_metrics.py
uv run pytest tests/test_dsight_sglang_log_metrics.py
uv run --with websockets python tests/dsight_log_metrics_check.py --port 9222 --out /tmp/dsight-log-metrics-check

The browser check requires a Chrome instance with remote debugging enabled. It builds synthetic source fixtures and checks paired visibility, zoomed summaries, changing/missing/conflicting limits, source evidence and narrow-screen layout.