Skip to content

DSight storage and query interface

DSight writes trace-data.sqlite for indexed local queries and an HTML catalog with compressed detail files for static browser delivery. Both contain the same normalized evidence. The browser never downloads or opens the SQLite database; it needs only index.html and the complete detail/ directory. Keep manifest.json with a published report to identify its generation and files.

See DSight for input discovery, generation commands, source coverage, and timing definitions. --single-file retains the embedded HTML and trace-data.json.gz format for direct file:// viewing.

SQLite schema, version 1

PRAGMA user_version is 1. This physical schema version is independent of the normalized JSON contract, srtctl-trace/1. Unsupported versions are rejected; rebuild from preserved inputs instead of modifying the cache in place.

Table Columns Key and meaning
catalog json TEXT NOT NULL, generation TEXT NOT NULL One row containing normalized metadata and a random generation ID. Requests, lifecycle, sessions, workers, batch observations, unmatched spans, source inventory, warnings and coverage remain in this JSON. Profile events/CPU samples and metric points are empty arrays; profile/sample counts and name/stack dictionaries remain.
events profile INTEGER, ordinal INTEGER, bucket INTEGER, start REAL, end REAL, name INTEGER, tid TEXT, source_row INTEGER PRIMARY KEY (profile, ordinal) WITHOUT ROWID. One imported NVTX interval. ordinal is its position in the normalized profile; name indexes that profile's name dictionary. tid preserves the global thread ID as text; source_row retains the original export row ID.
durations profile INTEGER, bucket INTEGER, maximum REAL PRIMARY KEY (profile, bucket) WITHOUT ROWID. Maximum observed interval duration in each duration class. Derived from events; no extra evidence is inferred.
samples profile INTEGER, ordinal INTEGER, time REAL, payload TEXT PRIMARY KEY (profile, ordinal) WITHOUT ROWID. CPU sample as JSON [time, global_tid, stack_index, source_row]; the stack and symbol dictionaries are in the profile's catalog entry.
points series INTEGER, ordinal INTEGER, time REAL, payload TEXT PRIMARY KEY (series, ordinal) WITHOUT ROWID. Metric point as JSON [time, value, source_id, source_row]. Series identity, labels, units, sample/setting semantics and configuration evidence remain in the catalog. Null values and duplicate-timestamp evidence are preserved.

Indexes are event_window(profile, bucket, start), sample_window(profile, time) and point_window(series, time). Profile and series IDs refer to their catalog entries; ordinal is a zero-based array position, not a source row ID. Source IDs refer to the catalog's source inventory. NVTX source identity comes from the profile's evidence_source.

Times remain the normalized floating-point seconds relative to catalog.meta.origin_ns, which is an exact string. The database does not recover raw nanoseconds discarded by an earlier import or correct clock skew. Negative setting timestamps, interval boundaries and source precision remain unchanged.

For NVTX, bucket = ceil(log2(max(end - start, 1e-9))). An overlap query uses each class's recorded maximum to bound both ends of an indexed start-time scan, then applies start <= to AND end >= from. This includes long intervals that started before the window without making every short interval scan reach back to the beginning of the capture. The lower bound is rounded outward with nextafter so floating-point subtraction cannot omit a boundary event. Results sort by (start, profile, source_row) before pagination; counts include all matching imported rows.

The cache opens with SQLite mode=ro. Query scratch tables use temporary storage. Readers check the generation ID before indexed queries and reject a replaced generation rather than combining old catalog metadata with new rows. Reopen the dataset after a rebuild. MCP detects replacement from file identity and reloads. The builder stages the database, detail files, HTML and manifest before replacing an existing managed report; failed builds retain the preceding generation.

Local query interface

The existing CLI, Python and MCP interfaces now accept either a report directory or an explicit trace-data.sqlite path. Directory lookup prefers SQLite; explicit legacy JSON/gzip paths still work. Opening SQLite reads the catalog, without loading the event, sample or point tables into Python lists.

uv run srtctl dsight query ./report summary
uv run srtctl dsight query ./report profiles --worker decode-0
uv run srtctl dsight query ./report nsys --profile 3 --from 600 --to 601 --limit 100
uv run srtctl dsight query ./report/trace-data.sqlite metrics \
  --name "<recorded_metric_name>" --from 600 --to 601 --points

Choose profile IDs and metric names from the report's catalogs; the values above are examples. CLI --from/--to map to Python/MCP start/end:

from srtctl.dsight.query import TraceDataset

trace = TraceDataset.from_path("./report")
profiles = trace.query("profiles", worker="decode-0")
page = trace.query("nsys", profile=profiles["items"][0]["id"],
                   start=600, end=601, offset=0, limit=100)

MCP exposes query_trace(dataset, kind="summary", **filters) with the same parameters and response shapes. The path belongs to the MCP server's filesystem. It is read-only and does not build reports or start profiling.

Query kinds Data access and filters
nsys Indexed interval overlap; profile, worker, recorded rank, case-insensitive name substring. Returns exact total, bounded rows, provenance and partial when any selected profile was truncated at import.
cpu Indexed inclusive sample window; profile, worker, recorded rank. Aggregates inclusive symbol hits, counting a symbol once per sample; returns total_samples and sorted hotspots.
metrics Indexed inclusive point window for selected series; worker, recorded rank, exact metric name. Streams aggregate statistics and optionally returns raw points.
profiles Catalog metadata and imported counts; profile, worker, recorded rank.
summary, requests, request, lifecycle, iterations, server_spans, sources Existing catalog queries. Request filters include session, agent, worker, search, min_ttft_ms; request and lifecycle require request_id. These collections have not been moved into relational tables.

All ranges require 0 <= start < end <= duration and default to the whole run. Intervals overlap inclusively; point/sample selection includes both endpoints. List responses contain total, items, offset, limit and range. Offset is nonnegative, limit defaults to 100 and is capped at 1,000; zero returns counts. For metrics, pagination selects series. points=True includes at most 1,000 raw points per series and reports points_total/points_truncated. Statistics use all non-null samples in the selected window. Settings additionally return every point at the latest timestamp strictly preceding the window in carried_setting; conflicting settings are not resolved by arbitrarily selecting one row.

To inspect the physical schema directly without changing it:

import sqlite3
from pathlib import Path

with sqlite3.connect(Path("report/trace-data.sqlite").resolve().as_uri() + "?mode=ro", uri=True) as db:
    print(db.execute("PRAGMA user_version").fetchone())
    print(db.execute("SELECT name, sql FROM sqlite_master WHERE type IN ('table', 'index')").fetchall())

Prefer TraceDataset.query for evidence queries; it owns dictionary resolution, interval indexing, source references, validation and response compatibility.

Static browser delivery and API

index.html embeds UI assets plus compressed catalog metadata, a 512-bin NVTX density overview per profile, and indexes of exact detail shards. Each shard has at most 8,192 rows and targets at most 512 KiB of decoded JSON (a single oversized row is kept intact). Names are the SHA-256 of compressed content. Descriptors retain compressed/decoded byte sizes, both hashes, row count and [min_start, min_end, max_start, max_end]. Point bounds repeat their timestamps.

The browser fetches selected metric families/windows, CPU samples and NVTX detail as needed. Whole selected shards can contribute exact NVTX counts without downloading their rows; boundary/name-filtered shards are decoded to check exact overlap. A short window renders exact ranges when candidate rows fit the 20,000 row display budget. Broader windows show an explicitly labeled density overview: overlapping interval counts, not CPU utilization or per-request execution time. Exact paginated queries remain available regardless of that display budget.

The shared LRU cache allows 32 MiB of decoded JSON source bytes. This is a cache accounting limit, not a bound on total JavaScript heap, chart arrays or an explicit full-window metric query. New views cancel stale rendering fetches. Missing or malformed shards produce errors; secure contexts (HTTPS/localhost) also verify decoded SHA-256. Static hosts may serve gzip files with or without Content-Encoding: gzip.

Browser API 3.1 makes Nsight and CPU detail methods awaitable for progressive reports. Use await with either delivery format:

const x = window.traceExplorer;
const profiles = x.listProfiles();                // synchronous metadata
const page = await x.queryNsys({profile: profiles[0].id, from: 600, to: 601, limit: 100});
await x.inspectNsys({profile: profiles[0].id, from: 600, to: 601});
await x.whenDetailsReady();                        // current rendered Nsight view
await x.whenMetricsReady();                        // current visible metric charts
const samples = await x.queryCpu({from: 600, to: 601});
const metrics = await x.queryMetrics({from: 600, to: 601, points: true});
const exported = await x.exportSelection();
const cache = x.detailDataStatus();

Browser queries keep their existing response shapes and selected-profile scope; they are not a SQL endpoint. queryNsys preserves profile source order, while local multi-profile queries sort by time/profile/source row. Catalog, request, lifecycle and batch queries remain synchronous. Query data is exact within the imported evidence; chart density/downsampling never substitutes for query rows.

Measured HTML loading impact

Measured on 2026-09-29 with unmodified main e7e8d007 and this change's renderer 9dd5fa98, using identical normalized inputs for each pair. The inputs are two preserved TokenSpeed inference captures; no inference benchmark was rerun.

Capture Client requests Imported NVTX rows Metric points Main HTML Progressive HTML
Before capacity fix 17 66,710 8,012,886 114.107 MB 1.805 MB
After capacity fix 406 4,162,788 7,616,598 185.104 MB 3.988 MB

MB uses decimal bytes. The HTML reduction is 98.4% and 97.8%, respectively; exact detail remains in adjacent files and the local SQLite cache.

Capture Main ready Progressive ready Main with visible metrics Progressive with visible metrics
Before capacity fix 560 ms 152 ms 633 ms 222 ms
After capacity fix 5,132 ms 328 ms 5,198 ms 388 ms

These are medians of three cache-disabled HTTP navigations per renderer in isolated headless Chrome 154.0.8037.57, Linux x86-64, Intel Core Ultra 9 285K (24 logical CPUs). Reports were served from localhost by Python's ThreadingHTTPServer; filesystem caches were not flushed and the browser process was reused. Ready measures navigation start to trace-explorer:ready, including HTML transfer, parse/decompression and initial UI rendering. The visible-metrics measurement additionally awaits whenMetricsReady(), current Nsight detail when available, and two animation frames; it includes browser-driver observation overhead. It does not load every metric family or every Nsight row.

One additional progressive navigation per capture, throttled to 10 Mbit/s and 20 ms latency in Chrome, measured 1.63 s / 3.53 s to ready and 2.09 s / 3.99 s with visible metrics. These controlled local measurements are not authenticated remote-host timings or a guarantee for other captures. The larger catalog is still about 4 MB, so initial transfer remains material on a constrained connection.

Both pages also passed a selected-window query check against SQLite: the tested one-second windows contained 44 and 3,009 ranges. Browser page counts and source row IDs matched. End-to-end window selection plus exact rendering took 34 ms and 128 ms on localhost in those single checks (after resetting throttling).

Reproduce the browser measurement with two report directories containing the same evidence and the respective HTML renderers:

uv run --with websockets python tests/dsight_bundle_check.py \
  --out /tmp/dsight-browser-check \
  --baseline /path/to/main-report --report /path/to/progressive-report --repeats 3

The harness starts an isolated Chrome and local HTTP server, saves navigation timings, resource sizes and screenshots, then compares exact browser/SQLite results. It also checks synthetic crossing windows, pagination, cancellation, cache accounting, exports, CPU samples, carried/conflicting settings and saved views. --dataset /path/to/local-report supplies the query cache when --report is a browser-only publication. Without report arguments it runs the synthetic regression fixture alone. A Chrome/Chromium executable must be on PATH.

Limits

The importer still materializes normalized inputs during a build. The default 250,000-event limit per Nsight report and its partial-import warnings are unchanged. Request/lifecycle catalogs and shard indexes still grow with the capture. SQLite and browser shards duplicate large arrays on disk; this trades build time and storage for faster reopening and lower initial browser work. Large full-range queries can still require substantial I/O, aggregation or memory. Original Nsight exports remain the source for excluded NVTX categories, CUDA kernel timing and evidence not imported into DSight.