Supported Models#

Supported checkpoint IDs are listed below. See the support matrix for platform coverage and the examples for workflows.

Support Policy#

TensorRT Edge-LLM supports the checkpoint IDs listed below. Dense LLM families include official dense checkpoints below 30B parameters. Larger dense checkpoints and non-dense variants require case-by-case validation. MoE, multimodal, audio, TTS, omni, EAGLE3, DFlash, and JetSpec support is limited to the listed rows.

The model coverage list is not comprehensive, and not every listed checkpoint has been fully verified on every supported platform and precision. If a listed model does not export, build, or run correctly, please report an issue with the checkpoint ID, precision, platform, and command line used.

The model class names were checked against the upstream Transformers model source tree. Checkpoint IDs are linked to their Hugging Face pages and grouped into original checkpoints and quantized checkpoints. All links below are public or gated; run hf auth login after accepting the provider’s terms for a gated checkpoint.

Precision Notes#

  • Dense precision set: FP16/BF16 checkpoints, ModelOpt FP8/MXFP8/FP4/NVFP4/INT4 AWQ/INT8 SmoothQuant checkpoints, and INT4 GPTQ checkpoints. INT8 GPTQ is not supported.

  • Jetson Orin supports FP16, INT8, and INT4 runtime precision in the supported JetPack configurations. Do not select FP8, MXFP8, FP4, or NVFP4 checkpoints for Orin.

  • For INT4 engine builds on Jetson Orin devices with less system memory, such as Jetson Orin Nano, pass --externalize-weights int4_ffn for dense checkpoints or --externalize-weights int4_ffn int4_moe for MoE checkpoints to reduce engine build memory.

  • For FP16/BF16 source checkpoints, use the Quantization script to create a unified quantized checkpoint for tensorrt_edgellm, then export the generated checkpoint.

  • FP8 KV cache is detected automatically from checkpoint metadata by tensorrt_edgellm.

  • tensorrt-edgellm-export exports visual encoders. Use tensorrt-edgellm-quantize llm --visual_quantization fp8 before export when FP8 visual weights are required.

  • MXFP8 and FP4/NVFP4 require Blackwell-class hardware for runtime execution.

  • For platform-specific NVFP4 MoE layouts, follow Export NVFP4 MoE for SM12x before building the engine.

Text Generation#

Llama 3.x

model_type: llama; text input to text output.

Original:

Pre-quantized:

Qwen2 / Qwen2.5

model_type: qwen2; text input to text output.

Original:

Pre-quantized:

Qwen3

model_type: qwen3 or qwen3_moe; text input to text output.

Original:

Pre-quantized:

HunYuan V1 Dense

model_type: hunyuan_v1_dense; text input to text output. These checkpoints apply per-head QK RMSNorm after RoPE and use the DynamicNTKAlpha RoPE variant (rope_scaling.type == "dynamic" with alpha); both are handled automatically by export and runtime.

Original:

Qwen vision-language families

model_type: qwen2_5_vl or qwen3_vl; text and image/video input to text output.

Qwen2.5-VL:

Qwen3-VL and compatible checkpoints:

Qwen3.5 / Qwen3.6 / Qwen3.8

model_type: qwen3_5 or qwen3_5_moe; text-only checkpoints accept text, while multimodal checkpoints accept text and image/video. Both produce text.

InternVL and Phi-4

InternVL uses model_type: internvl_chat or internvl; Phi-4 Multimodal uses phi4mm or phi4_multimodal. Both accept text and image input and produce text; Phi-4 also accepts audio.

InternVL:

Phi-4 Multimodal:

Nemotron-3 / 3.5

Nemotron-H language checkpoints use model_type: nemotron_h and accept text input. Nemotron Omni uses NemotronH_Nano_Omni_Reasoning_V3 and accepts text, image/video, and audio. Both produce text.

Nemotron-H:

Nemotron Omni:

Gemma4

model_type: gemma4 or gemma4_unified; accepted modalities depend on the checkpoint and all variants produce text. Paired gemma4_assistant or gemma4_unified_assistant checkpoints provide MTP.

Text, image, and audio input:

Text and image input:

MTP assistants:

DiffusionGemma

model_type: diffusion_gemma or diffusiongemma; text and image input to text output through block-diffusion decoding.

Muse-Glimmer

model_type: muse_glimmer; text, image, or video input to text output.

The FP16 checkpoint supports text, image, and video input; the NVFP4 checkpoint is text-only. Compatible DFlash and DFlash2 drafts are listed under Speculative Draft Checkpoints.

Speech Recognition#

Speech Generation#

Action and Multimodal Reasoning#

Speculative Draft Checkpoints#

EAGLE3 Draft Models#

DFlash Draft Models#

DFlash2 Draft Models#

DFlash2 uses the public DFlash engine/runtime mode. The draft checkpoint selects the versioned contract, while the runtime may choose any block size covered by the engine profile. DFlash2 does not support DDTree.

Draft checkpoint

Base model

z-lab/Qwen3.8-27B-DFlash2

Qwen/Qwen3.8-27B, including matched NVFP4 or INT4 quantized checkpoints

incoai/Muse-Glimmer-30B-DFlash2

meta-models/Muse-Glimmer-30B or RadixArk/Muse-Glimmer-NVFP4

DSpark Draft Models#

The Nemotron-3.5-Lightning DSpark draft differs from the DeepSpec block7 drafts above: it uses a block size of 8, sliding-window attention (1024) with a learned per-head attention sink, and causal proposal attention. The sliding window is baked into the draft engine as the contiguous-query XQA variant. The Nemotron draft supports both chain decoding and greedy DSpark DDTree; tree mode reconstructs path-dependent recurrent state. The published checkpoint uses sample_from_anchor=false: slot 0 is anchor-only and proposals start at slot 1. The runtime also supports the older anchor-sampled layout. Tree mode uses --specDraftTopK > 1 with a tree-capable base.

JetSpec Draft Models#

JetSpec draft checkpoints are detected by jetspec_config in config.json and exported with DFlashDraftModel using causal proposal attention. The validated runtime path is branching tree verification: export the base with --jetspec-tree-base --jetspec-draft-dir <draft_checkpoint>, export the draft with --jetspec-draft --jetspec-draft-dir <draft_checkpoint>, and run with --specDraftTopK > 1. --jetspecBlockSize and --dflashBlockSize configure the same cached-draft proposal block size; the JetSpec spelling is provided for clarity in JetSpec command lines.

For the listed Qwen3 pair, disable thinking mode in the input JSON when evaluating accuracy, acceptance rate, or throughput.

Draft checkpoint

Base model

JetSpec/jetspec-qwen3-8b

Qwen/Qwen3-8B