nemotron model family
Tasks
Supported architectures and task heads
The Hugging Face values below are copied from checkpoint metadata at the recorded revision. The TRTMC task contract comes from the exact E2E recipe. Architecture identity selects or describes a source graph; it does not imply that TRTMC reproduces every Hugging Face head with the same model_type. For example, an encoder recipe may consume the base model and intentionally return hidden states instead of the checkpoint's pretraining or classification logits.
HF model_type | HF architecture / pipeline class | TRTMC task contract | Exact recipe profiles |
|---|---|---|---|
nemotron | NemotronForCausalLMSource: config.architecturesMetadata: config.json at 05c6ef6d2dda | text_generation_causal | nemotron-hindi-4bnemotron-mini-4b |
Declared recipes
Each row comes from a model-owned E2E manifest. It declares a test recipe; it is not a live pass receipt for every hardware target.
| Recipe | Exact Hugging Face checkpoint | Task | Build configuration | Declared E2E cases | Manifest |
|---|---|---|---|---|---|
nemotron-hindi-4b | nvidia/Nemotron-4-Mini-Hindi-4B-Base | bf16; single device | nemotron-hindi-4b | tests/e2e/models/nemotron/manifests/nemotron-hindi-4b.json | |
nemotron-mini-4b | nvidia/Nemotron-Mini-4B-Instruct | fp16; single device | nemotron-mini-4b | tests/e2e/models/nemotron/manifests/nemotron-mini-4b.json |
Family-owned configuration
This family does not declare a family-owned --set namespace. Use the explicit CLI options shown below and the shared configuration namespaces documented in Configure Runtime Behavior.
Family-specific CLI contracts
Inputs and options below are filtered by the declared E2E task and the methods and configuration fields used by that family's native runtime implementation. The global CLI parser accepts a wider union of flags; flags absent here are not declared for this family.
trtmc run
Generate text from a text prompt.
Declared recipes: nemotron-hindi-4b, nemotron-mini-4b
trtmc run <bundle.bundle> --prompt "<text>" [generation options]| Supported input or option | Requirement | Runtime behavior |
|---|---|---|
--prompt <TEXT> | Required | Text input for this causal language-model recipe. |
--max-new-tokens <N> | Optional | Limit the number of generated tokens or audio frames. |
--temperature <F> | Optional | Set sampling temperature. |
--top-k <N> | Optional | Restrict sampling to the top K tokens. |
--top-p <F> | Optional | Enable nucleus sampling at probability P. |
--min-p <F> | Optional | Filter tokens below min-p times the maximum probability. |
--seed <N> | Optional | Set the sampling or diffusion seed. |
--chat-template | Optional | Apply the chat template packaged with the bundle. |
--no-thinking | Optional | Disable a supported reasoning or thinking mode. |
--greedy | Optional | Select deterministic greedy decoding by setting temperature to zero. |
--num-samples <N> | Optional | Run N independent text generations. |
--output <PATH> | Optional | Write generated samples as JSON Lines. |
--benchmark <N> | Optional | Run N timed generation iterations. |
--warmup <N> | Optional with --benchmark | Warm-up iterations before generation timing. |
Code basis
Runtime provider: nemotron
include/trtmc/pipeline.hsrc/cli/main.cppsrc/runtime/models/nemotron/pipeline.cppsrc/runtime/models/nemotron/sampler.cpptests/e2e/models/nemotron/manifests/nemotron-hindi-4b.jsontests/e2e/models/nemotron/manifests/nemotron-mini-4b.json
Bundle lifecycle commands
These commands apply to every declared recipe in this family; they do not add task inputs or model capabilities.
trtmc build
Build one exact checkpoint into a TensorRT-Model-Connect bundle.
trtmc build <hf-id> --output <bundle.bundle>trtmc inspect
Inspect bundle metadata, runtime identity, and packaged sections.
trtmc inspect <bundle.bundle>See the CLI Reference for all options and limitations.