Skip to main content

← All model recipe tasks

codegen model family

Tasks

Supported architectures and task heads

The Hugging Face values below are copied from checkpoint metadata at the recorded revision. The TRTMC task contract comes from the exact E2E recipe. Architecture identity selects or describes a source graph; it does not imply that TRTMC reproduces every Hugging Face head with the same model_type. For example, an encoder recipe may consume the base model and intentionally return hidden states instead of the checkpoint's pretraining or classification logits.

HF model_typeHF architecture / pipeline classTRTMC task contractExact recipe profiles
codegenCodeGenForCausalLM
Source: config.architectures
Metadata: config.json at d9107f71cca4
text_generation_causal
codegen-350m
codegen-350m-tp4

Declared recipes

Each row comes from a model-owned E2E manifest. It declares a test recipe; it is not a live pass receipt for every hardware target.

RecipeExact Hugging Face checkpointTaskBuild configurationDeclared E2E casesManifest
codegen-350mSalesforce/codegen-350M-monofp16; single devicecodegen-350mtests/e2e/models/codegen/manifests/codegen-350m.json
codegen-350m-tp4Salesforce/codegen-350M-monofamily default; TP4codegen-350m-tp4tests/e2e/models/codegen/manifests/codegen-350m-tp4.json

Family-owned configuration

This family does not declare a family-owned --set namespace. Use the explicit CLI options shown below and the shared configuration namespaces documented in Configure Runtime Behavior.

Family-specific CLI contracts

Inputs and options below are filtered by the declared E2E task and the methods and configuration fields used by that family's native runtime implementation. The global CLI parser accepts a wider union of flags; flags absent here are not declared for this family.

trtmc run

Generate text from a text prompt.

Declared recipes: codegen-350m, codegen-350m-tp4

trtmc run <bundle.bundle> --prompt "<text>" [generation options]
Supported input or optionRequirementRuntime behavior
--prompt <TEXT>RequiredText input for this causal language-model recipe.
--max-new-tokens <N>OptionalLimit the number of generated tokens or audio frames.
--temperature <F>OptionalSet sampling temperature.
--top-k <N>OptionalRestrict sampling to the top K tokens.
--top-p <F>OptionalEnable nucleus sampling at probability P.
--min-p <F>OptionalFilter tokens below min-p times the maximum probability.
--seed <N>OptionalSet the sampling or diffusion seed.
--chat-templateOptionalApply the chat template packaged with the bundle.
--no-thinkingOptionalDisable a supported reasoning or thinking mode.
--greedyOptionalSelect deterministic greedy decoding by setting temperature to zero.
--num-samples <N>OptionalRun N independent text generations.
--output <PATH>OptionalWrite generated samples as JSON Lines.
--benchmark <N>OptionalRun N timed generation iterations.
--warmup <N>Optional with --benchmarkWarm-up iterations before generation timing.
Code basis

Runtime provider: codegen

  • include/trtmc/pipeline.h
  • src/cli/main.cpp
  • src/runtime/models/codegen/pipeline.cpp
  • src/runtime/models/codegen/sampler.cpp
  • tests/e2e/models/codegen/manifests/codegen-350m-tp4.json
  • tests/e2e/models/codegen/manifests/codegen-350m.json

Bundle lifecycle commands

These commands apply to every declared recipe in this family; they do not add task inputs or model capabilities.

trtmc build

Build one exact checkpoint into a TensorRT-Model-Connect bundle.

trtmc build <hf-id> --output <bundle.bundle>

trtmc inspect

Inspect bundle metadata, runtime identity, and packaged sections.

trtmc inspect <bundle.bundle>

See the CLI Reference for all options and limitations.