Skip to main content

Validate a Model Contribution

Use this workflow after Add a Model Family, or after changing an existing family's build, runtime, dependency, or oracle.

Identify the ownership unit​

A family contribution is one vertical slice:

families/<family>/
support.py
model.py
requirements.txt # optional
runtime/CMakeLists.txt
runtime/*.cpp
tests/test_*.py
tests/manifests/*.json
tests/thresholds/*.json # optional numeric overrides

Record the family, exact model ID and immutable revision, task, testcase, precision, quantization, tensor/context parallel sizes, shape bounds, required assets/dependencies, native targets, GPU count, and oracle before testing.

1. Validate ownership and impact​

python3 -m tools.model_ci validate
python3 tools/test_impact.py --validate
python3 tools/test_impact.py --base github/main --head HEAD

The impact report should select the owning family. Production source must not depend on a sibling family, and adding a normal family must not modify a central registry or source list.

2. Run source and unit contracts​

python3 -m pytest core/builder/tests tools/tests -q
python3 -m pytest families/<family>/tests -m 'not gpu and not trt and not e2e' -q

Build and run the exact family-owned native targets when present:

cmake -S . -B build -DTRTMC_BUILD_TESTS=ON
cmake --build build --target trtmc_model_<family>
ctest --test-dir build --output-on-failure

Passing shared architecture tests does not replace family-specific proof.

3. Build and inspect a representative bundle​

python -m tensorrt_model_connect build Qwen/Qwen3-0.6B \
--precision fp16 \
--output /tmp/qwen3-0.6b.bundle
trtmc inspect /tmp/qwen3-0.6b.bundle

Verify the exact family, task, backend, and required family-owned sections. Inspection proves container construction, not inference parity.

4. Run the declared family E2E​

Family test_e2e.py files accept an explicit selection and require the native binary/runtime root through the environment:

TRTMC_BINARY="$PWD/build/apps/cli/trtmc" \
TRTMC_RUNTIME_ROOT="$PWD/build/install/lib" \
python3 -m pytest families/qwen/tests/test_e2e.py \
--e2e-testcase qwen3-0.6b-fp16 \
-q -x

The selected test downloads or opens its declared checkpoint, builds through the public Python API, loads exactly the owning family DSO, invokes the public Task API, and applies its own oracle. Tensor-parallel cases additionally need the declared GPU count and mpirun.

Confirm that the manifest has a meaningful task/testcase, exact inputs, premerge selection where intended, and thresholds that reject adversarial or known-wrong outputs. Never weaken a criterion to pass CI.

5. Save and read correctness evidence​

Set an artifact directory when running a selected family E2E to retain the inputs, native and reference outputs, original assertion expressions and evaluated values, stage timings, and reproduction commands:

TRTMC_E2E_ARTIFACT_DIR=/tmp/trtmc-e2e \
TRTMC_BINARY="$PWD/build/apps/cli/trtmc" \
TRTMC_RUNTIME_ROOT="$PWD/build/install/lib" \
python3 -m pytest families/qwen/tests/test_e2e.py \
--e2e-testcase qwen3-0.6b-fp16 -q

python3 -m tools.e2e_report /tmp/trtmc-e2e \
-o /tmp/trtmc-correctness.html

Open the HTML directly in a browser. Each testcase also writes its own evidence/<family>/<case>/evidence.json and report.html; failure paths retain the observations produced before the failure. Use a fresh artifact directory for each campaign. Reusing a directory replaces earlier evidence for the selected family and case, while untouched cases remain from their original runs.

The report records existing assertions; it does not replace a family's oracle or change thresholds. A contract-only check is identified as such instead of claiming an official-reference comparison. Model-specific visualizations remain in the owning family's tests. Detailed arrays and original files are retained alongside bounded, embedded media; any omitted or truncated evidence is marked. The standalone report embeds up to 32 MiB per media file and 256 MiB across the report. Full-size raw data stays in the testcase evidence directory.

When a library enforces the comparison itself, record every original check without adding duplicate pytest assertions. Use independent_reference for native/reference output comparisons and contract for counts or finiteness:

record_evidence("reference_comparison", {
"label": "original request",
"scope": "independent_reference",
"enforced": True,
"native": native_path, # Path to an existing output file
"reference": reference_path, # Path to a distinct retained reference file
"checks": [{
"name": "similarity", "label": "Similarity",
"scope": "independent_reference",
"actual": metrics["similarity"], "operator": ">=",
"expected": thresholds["similarity_min"],
"passed": checks["similarity"],
}], # Include all checks from the enforced library comparison.
})

Emit each request's comparison separately, retaining its measured values, limits, and original verdicts. Supported operators are ==, >=, and <=. The report requires complete, consistent check rows and a passed testcase; passed alone, a reference file alone, or only contract checks cannot establish reference verification. Record diagnostics without replacing library failures.

Distinguish a testcase's actual execution result from certification of its whole family or pipeline. A family can fail while some of its cases pass. Also distinguish correctness from the latency and throughput measurements in trtmc-bench reports. Compare the same checkpoint, inputs, seed or initial latents, precision, and runtime configuration before interpreting a difference.

6. Report evidence by level​

LevelWhat it establishes
Repository-consistentOwnership and impact validators accept the tree.
Unit-testedFocused builder, runtime, tool, and family contracts pass.
Inference-testedThe exact bundle runs its declared Task on compatible hardware.
Parity-qualifiedRetained comparison artifacts pass the intended official-reference contract.
Performance-qualifiedExact-hardware results retain inputs, warmups, repetitions, baseline, and raw measurements.

State exactly which revision, checkpoint, hardware, commands, and cases ran, plus unverified paths. For the full CI path and one-shot protected-CI label, follow Contributing.