Skip to main content

Benchmarking reference

The benchmark application sits above TRTMC's public build and task APIs. It does not register model families or import family implementation code.

Single-model benchmark​

Install TRTMC, then select an installed runtime root containing the core runtime, TensorRT backend, and required family DSOs:

trtmc-bench list models
trtmc-bench run \
--model distilgpt2 \
--runtime-root /opt/trtmc/lib \
-o results/distilgpt2

trtmc-bench reads family-owned manifests under families/*/tests/manifests/. Missing bundles are built through the public builder and cached; pass --no-build when every selected bundle must already exist.

Run a checked-in multi-model specification with:

trtmc-bench run apps/benchmark/example.yaml -o results/example

Use trtmc-bench --help and subcommand help as the authoritative option list for the installed version.

Timing boundary​

Candidate timing measures the public family task call. Bundle construction, process startup, task loading, warmup, telemetry, and report generation are outside that measurement. Asset loading is also excluded unless a case explicitly includes it.

The native worker stops the clock before formatting an observation or destroying the previous iteration's result. New receipts explicitly include observation_serialization_included: false. Earlier worker revisions included observation formatting in the measured duration; account for that boundary change when comparing historical measurements. No comparison threshold changes with this correction.

Generated media is retained automatically for semantic tasks. Each measured audio output is an interleaved FLOAT32 WAV; images and ordered video frames use the existing CLI PNG encoding. Consumed conditioning images are also retained. The JSON observations reference these files relative to the case directory, so keep that directory with its report when moving or archiving results. Warmup does not write media, and the summary references the final measured result. Legal empty audio is explicitly marked empty, without a fabricated waveform.

Speech-dialogue observations retain the consumed input as a FLOAT32 WAV and separate WAVs for audio events. event_audio_artifacts follows the unchanged events array: null entries have no audio file. The HTML shows input audio, effective system prompt, text/tool events and per-event players with epoch and sequence numbers. It does not splice events or epochs into a fabricated timeline.

WAV and PNG serialization is outside the call timer. Streaming TTS must copy borrowed callback samples while they are valid; that copy is inside the public call and is declared by streaming_pcm_copy_included: true. Account for this receiving cost when comparing a streaming reference. PNGs are visual previews, not a substitute for the family's original floating-point correctness checks. Dialogue input buffers remain alive through observation so the retained input is the decoded data actually supplied, not a later reread of the source file. Their release, like result release, is outside the measured call; the existing asset-loading option still controls whether file decoding is timed. For older dialogue workers with timed asset loading, decoded-input release was inside the timer; account for this boundary change in historical comparisons.

The HTML report displays recorded text, playable audio, and image/frame previews next to the input prompt. Image editing also shows the images actually supplied to the task. Video frames retain their recorded order and timestamps; no frame rate is inferred when a timeline is absent. Missing or unsafe media references are shown as unavailable rather than embedded from outside the case directory.

The release suite defaults to three warmups and ten measured iterations. A reference within five percent of candidate p50 is considered equivalent. Candidate or reference execution failures are operational failures, not slow performance results.

Internal release performance matrix​

This repository-only workflow is CI qualification. It is separate from the installed trtmc-bench user interface above and is not included in the wheel.

The matrix coordinates candidate and reference runs without introducing a second model registry:

export TRTMC_PERF_WORKER=/opt/trtmc/bin/trtmc_benchmark_worker
export TRTMC_PERF_RUNTIME_ROOT=/opt/trtmc/lib
export TRTMC_PERF_BUNDLE_CACHE=/data/trtmc-bundles

python3 -m qualification_tests.benchmark_qualification.performance check \
qualification_tests/benchmark_qualification/performance/config/release.yaml \
--environment qualification_tests/benchmark_qualification/performance/config/environments/gb300.yaml

python3 -m qualification_tests.benchmark_qualification.performance prepare \
qualification_tests/benchmark_qualification/performance/config/release.yaml \
--environment qualification_tests/benchmark_qualification/performance/config/environments/gb300.yaml \
--entry gpt2.generate \
--output artifacts/perf/bundle-preparation.json

python3 -m qualification_tests.benchmark_qualification.performance run \
qualification_tests/benchmark_qualification/performance/config/release.yaml \
--environment qualification_tests/benchmark_qualification/performance/config/environments/gb300.yaml \
--entry gpt2.generate

Preparation is deliberately separate and untimed. Resume or regenerate a report from stored observations with:

python3 -m qualification_tests.benchmark_qualification.performance resume artifacts/perf/<run-directory>
python3 -m qualification_tests.benchmark_qualification.performance report artifacts/perf/<run-directory>

The release YAML owns model, testcase, operation, measurement, reference, and comparison semantics. Machine-specific paths belong in the environment YAML. Reference implementations run in separate processes so their dependencies do not enter the candidate worker or shared runtime.

Adding coverage​

  • Add a weight or profile to the owning family's test manifest catalog.
  • Add a matrix entry that names the family, model, operation, workload, and reference runner.
  • Add benchmark code only when a genuinely new public task interface needs a task adapter.

Retain raw observations and reports together with the commit, model revision, runtime root, target GPU, TensorRT version, warmup count, and iteration count. Correctness validation should precede performance comparison.

Semantic Task SDK integration​

Migrated bundles select the C Task SDK before execution; non-migrated bundles retain their existing path. A failed C call is never retried through an old interface. The worker currently has semantic routes for prompt continuation, conditional generation, corrupted-text reconstruction, summarization, image/text generation, point, quantile or joint forecasting, and the following media, perception and feature workloads. An SDK Task not listed here needs its own explicit benchmark inputs; a shared operation name does not imply support.

OperationSemantic Tasks
generate_audioTextToAudio, TextToSpeech, StreamingTextToSpeech
transcribeSpeechTranscription, SpeechTranslation, StreamingSpeechTranscription
speakSpeechToSpeechResponse
generate_imageTextToImage, ImagesTextToImageEdit, BatchTextToImage, TextToVideo, ImageTextActionToVideo
classifyImageToClassScores
extract_featuresImageToTokenFeatures, ImageToPooledFeatures, ImageToTokenAndPooledFeatures, ImageToSpatialFeatures
segmentImageToSemanticSegmentation, ImagePointsToMasks center helper
segment_promptedImagePointsToMasks, ImageTextToInstanceMasks
disparityStereoImagesToDisparity
encodeTextToPooledFeatures, TextToTokenFeatures
embedTextToEmbedding
rerankTextQueryDocumentsToRelevance
controlImageStateToActionChunk

For semantic Task manifests, explicitly specified generation controls retain their types; omitted controls use the family defaults. A testcase can supply additional family-declared keys through its config object. Flat and nested duplicate keys are errors, not an implicit override mechanism. Model data stays in typed request fields, never in Config. A quantile or channel axis is not a request batch. Joint forecasts retain both point and quantile outputs from one family evaluation.

Audio SDK inputs retain WAV sample rate and interleaved channels. Duration uses frames divided by sample rate; num_samples/output_samples remain scalar PCM counts, with explicit channels and output_frames. Transcription retains text, token IDs and the family's actual segment timestamps. The worker's existing max_new_tokens spelling maps to the declared max_output_tokens Config key for one-shot transcription/translation only; streaming ASR keeps max_new_tokens. language is the typed ASR source language or TTS language, and translation has a separate optional target_language. Absent languages use family defaults.

Streaming ASR measures a fresh stream for every invocation, including creation, chunk submission, finalization and release. Its separate first_partial_ms clock begins after stream creation. Packetization defaults to 160 ms and never splits an interleaved frame. Streaming TTS measures the direct synchronous callback API, including copies of borrowed PCM. Retained WAV files are written after that call, outside its timer; no extra worker or inference is added. An explicit streaming input must agree with the semantic Task. Failed or stopped calls are operational failures, never completed measurements. Native-batch audio and dialogue-session benchmarks are not included in these routes.

Image batches make one native batch call with ordered prompts and optional seeds/item_configs arrays. Their lengths must match the request count, and shared/item duplicate keys are rejected. Scalar replay uses initial_latents_path; this does not define a broadcast input for a batch. Video metrics count actual clips and all returned frames, including any conditioned_prefix_frames. Thus frames_per_s is returned-frame throughput, not newly predicted-frame throughput or an inferred playback FPS. Returned timestamps and the conditioned prefix remain explicit in the output. A worker-only completion has no produced media and fails this single-process benchmark instead of counting as an image or video.

Pooled and token features remain different workloads. The joint image feature Task returns both hidden and pooled arrays from one call. Classification keeps the original score representation, and document-list reranking preserves order; the list interface alone is not evidence of native GPU batching. SAM's segment helper uses a center foreground point and thresholds the first family-selected mask at zero; explicit points use segment_prompted, which keeps all masks. Empty masks remain empty. Confidence and predicted IoU are reported separately.

Disparity writes the final F32 map beside the result JSON after measurement, retaining dimensions and the left-grid x_left_minus_x_right convention. File flush errors fail the run. New semantic image/feature routes honor an explicit asset_loading_included=true; older legacy image/feature paths did not all honor that flag, so historical comparisons must account for this boundary. Stateless action chunks preserve schema, values and timing; they do not emulate a queue.

trtmc_dataset_benchmark preserves its separate, explicit benchmark workload: 12000 maximum new tokens, temperature/top-p 1, top-k 1, min-p 0, seed -1, chat formatting off, thinking on, boxed-answer stopping off and check interval 16. These defaults apply only to controls declared by the selected Task; an explicit unsupported flag fails even when its value equals a default. They are application workload choices, not shared runtime or QuickStart defaults. --set NAME=VALUE accepts additional declared scalar/list controls on semantic bundles, using the same text parser as the CLI. Each successful sample records the exact submitted_config; absent, family-computed values are not fabricated.

Use --task TASK_ID to select a bound prompt-only Task, including a secondary text capability of a multimodal bundle. Without it, the dataset benchmark uses a prompt-compatible primary Task, or the sole bound prompt-compatible Task. Multiple candidates require an explicit selection. Unknown, unbound or incomplete-input Tasks fail; a failed call never retries another Task or the existing family path. The result's task identifies the executed Task and bundle_task retains the bundle's primary identity. Existing bundle modes reject --task, just as they reject --set.

Translation passes the prompt as source text, leaving source and target languages absent so the family can use its declared bundle defaults. A family without an applicable default rejects the request; the benchmark does not invent a language. Tasks needing additional inputs, such as images or latent tensors, are not complete dataset-prompt requests and are not synthesized from the prompt.

For nonnegative base seeds, both dataset paths add seed_index when present, otherwise the sample's zero-based row index. Overflow fails instead of wrapping. Dataset timing ends before answer extraction and JSON output. Preserve the bundle, submitted parameters and software revision when comparing these runs.