Supported Models#
Supported checkpoint IDs are listed below. See the support matrix for platform coverage and the examples for workflows.
Support Policy#
TensorRT Edge-LLM supports the checkpoint IDs listed below. Dense LLM families include official dense checkpoints below 30B parameters. Larger dense checkpoints and non-dense variants require case-by-case validation. MoE, multimodal, audio, TTS, omni, EAGLE3, DFlash, and JetSpec support is limited to the listed rows.
The model coverage list is not comprehensive, and not every listed checkpoint has been fully verified on every supported platform and precision. If a listed model does not export, build, or run correctly, please report an issue with the checkpoint ID, precision, platform, and command line used.
The model class names were checked against the upstream Transformers model source tree. Checkpoint IDs are linked to their Hugging Face pages and grouped into original checkpoints and quantized checkpoints. All links below are public or gated; run hf auth login after accepting the provider’s terms for a gated checkpoint.
Precision Notes#
Dense precision set: FP16/BF16 checkpoints, ModelOpt FP8/MXFP8/FP4/NVFP4/INT4 AWQ/INT8 SmoothQuant checkpoints, and INT4 GPTQ checkpoints. INT8 GPTQ is not supported.
Jetson Orin supports FP16, INT8, and INT4 runtime precision in the supported JetPack configurations. Do not select FP8, MXFP8, FP4, or NVFP4 checkpoints for Orin.
For INT4 engine builds on Jetson Orin devices with less system memory, such as Jetson Orin Nano, pass
--externalize-weights int4_ffnfor dense checkpoints or--externalize-weights int4_ffn int4_moefor MoE checkpoints to reduce engine build memory.For FP16/BF16 source checkpoints, use the Quantization script to create a unified quantized checkpoint for
tensorrt_edgellm, then export the generated checkpoint.FP8 KV cache is detected automatically from checkpoint metadata by
tensorrt_edgellm.tensorrt-edgellm-exportexports visual encoders. Usetensorrt-edgellm-quantize llm --visual_quantization fp8before export when FP8 visual weights are required.MXFP8 and FP4/NVFP4 require Blackwell-class hardware for runtime execution.
For platform-specific NVFP4 MoE layouts, follow Export NVFP4 MoE for SM12x before building the engine.
Text Generation#
Llama 3.x
model_type: llama; text input to text output.
Original:
Pre-quantized:
Qwen2 / Qwen2.5
model_type: qwen2; text input to text output.
Original:
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B, deepseek-ai/DeepSeek-R1-Distill-Qwen-7B, deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Pre-quantized:
Qwen/Qwen2-0.5B-Instruct-AWQ, Qwen/Qwen2-0.5B-Instruct-GPTQ-Int4
Qwen/Qwen2-1.5B-Instruct-AWQ, Qwen/Qwen2-1.5B-Instruct-GPTQ-Int4
Qwen/Qwen2-7B-Instruct-AWQ, Qwen/Qwen2-7B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-0.5B-Instruct-AWQ, Qwen/Qwen2.5-0.5B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-1.5B-Instruct-AWQ, Qwen/Qwen2.5-1.5B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-3B-Instruct-AWQ, Qwen/Qwen2.5-3B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-7B-Instruct-AWQ, Qwen/Qwen2.5-7B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-14B-Instruct-AWQ, Qwen/Qwen2.5-14B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-0.5B-Instruct-AWQ, Qwen/Qwen2.5-Coder-0.5B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-1.5B-Instruct-AWQ, Qwen/Qwen2.5-Coder-1.5B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-3B-Instruct-AWQ, Qwen/Qwen2.5-Coder-3B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-7B-Instruct-AWQ, Qwen/Qwen2.5-Coder-7B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-14B-Instruct-AWQ, Qwen/Qwen2.5-Coder-14B-Instruct-GPTQ-Int4
Qwen3
model_type: qwen3 or qwen3_moe; text input to text output.
Original:
Pre-quantized:
HunYuan V1 Dense
model_type: hunyuan_v1_dense; text input to text output. These checkpoints
apply per-head QK RMSNorm after RoPE and use the
DynamicNTKAlpha RoPE variant (rope_scaling.type == "dynamic" with alpha);
both are handled automatically by export and runtime.
Original:
Qwen vision-language families
model_type: qwen2_5_vl or qwen3_vl; text and image/video input to text
output.
Qwen2.5-VL:
Qwen/Qwen2.5-VL-3B-Instruct-AWQ, Qwen/Qwen2.5-VL-7B-Instruct-AWQ
nvidia/Qwen2.5-VL-7B-Instruct-FP8, nvidia/Qwen2.5-VL-7B-Instruct-NVFP4
Qwen3-VL and compatible checkpoints:
Qwen3.5 / Qwen3.6 / Qwen3.8
model_type: qwen3_5 or qwen3_5_moe; text-only checkpoints accept text,
while multimodal checkpoints accept text and image/video. Both produce text.
InternVL and Phi-4
InternVL uses model_type: internvl_chat or internvl; Phi-4 Multimodal uses
phi4mm or phi4_multimodal. Both accept text and image input and produce
text; Phi-4 also accepts audio.
InternVL:
OpenGVLab/InternVL3_5-1B-HF, OpenGVLab/InternVL3_5-2B-HF, OpenGVLab/InternVL3_5-4B-HF
OpenGVLab/InternVL3-1B-AWQ, OpenGVLab/InternVL3-2B-AWQ, OpenGVLab/InternVL3-8B-AWQ, OpenGVLab/InternVL3-14B-AWQ
Phi-4 Multimodal:
Nemotron-3 / 3.5
Nemotron-H language checkpoints use model_type: nemotron_h and accept text
input. Nemotron Omni uses NemotronH_Nano_Omni_Reasoning_V3 and accepts text,
image/video, and audio. Both produce text.
Nemotron-H:
Nemotron Omni:
Gemma4
model_type: gemma4 or gemma4_unified; accepted modalities depend on the
checkpoint and all variants produce text. Paired gemma4_assistant or
gemma4_unified_assistant checkpoints provide MTP.
Text, image, and audio input:
Text and image input:
MTP assistants:
DiffusionGemma
model_type: diffusion_gemma or diffusiongemma; text and image input to text
output through block-diffusion decoding.
Muse-Glimmer
model_type: muse_glimmer; text, image, or video input to text output.
The FP16 checkpoint supports text, image, and video input; the NVFP4 checkpoint is text-only. Compatible DFlash and DFlash2 drafts are listed under Speculative Draft Checkpoints.
Speech Recognition#
Qwen3-ASR (
qwen3_asr): audio to transcript text. Qwen/Qwen3-ASR-0.6B, Qwen/Qwen3-ASR-1.7BNemotron-3.5-ASR (
nemotron3_5_asr): streaming audio to transcript text through its model-specific RNN-T runtime. nvidia/nemotron-3.5-asr-streaming-0.6b
Speech Generation#
Qwen3-TTS (
qwen3_tts): text, language, style, or reference-speech conditioning to speech. Qwen/Qwen3-TTS-12Hz-0.6B-CustomVoice, Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice, Qwen/Qwen3-TTS-12Hz-1.7B-VoiceDesign, Qwen/Qwen3-TTS-12Hz-0.6B-Base, Qwen/Qwen3-TTS-12Hz-1.7B-BaseQwen3-Omni (
qwen3_omni_moe): text, image/video, and audio to text and optional speech. Qwen/Qwen3-Omni-30B-A3B-Instruct
Action and Multimodal Reasoning#
Alpamayo (
alpamayo_r1): text and image/video to reasoning and a trajectory. nvidia/Alpamayo-R1-10BCosmos3-Edge (
cosmos3_edge): text and image/video to reasoning. nvidia/Cosmos3-Edgepi0.5 (model-specific exporter/runtime): camera observations and an instruction to a robot action chunk. lerobot/pi05_libero_base, lerobot/pi05_droid, lerobot/pi05_base (served under the
pi05_alohacontract)
Speculative Draft Checkpoints#
EAGLE3 Draft Models#
Draft checkpoint |
Base model |
|---|---|
DFlash Draft Models#
Draft checkpoint |
Base model |
|---|---|
z-lab/Qwen3.5-4B-DFlash (quantized checkpoint: |
Qwen3.5-4B-NVFP4 |
deepseek-ai/dflash_gemma4_12b_block7 or z-lab/gemma4-12B-it-DFlash |
|
DFlash2 Draft Models#
DFlash2 uses the public DFlash engine/runtime mode. The draft checkpoint selects the versioned contract, while the runtime may choose any block size covered by the engine profile. DFlash2 does not support DDTree.
Draft checkpoint |
Base model |
|---|---|
Qwen/Qwen3.8-27B, including matched NVFP4 or INT4 quantized checkpoints |
|
DSpark Draft Models#
Draft checkpoint |
Base model |
|---|---|
The Nemotron-3.5-Lightning DSpark draft differs from the DeepSpec block7 drafts
above: it uses a block size of 8, sliding-window attention (1024) with a learned
per-head attention sink, and causal proposal attention. The sliding window is
baked into the draft engine as the contiguous-query XQA variant. The Nemotron
draft supports both chain decoding and greedy DSpark DDTree; tree mode
reconstructs path-dependent recurrent state. The published checkpoint uses
sample_from_anchor=false: slot 0 is anchor-only and proposals start at slot 1.
The runtime also supports the older anchor-sampled layout. Tree mode uses
--specDraftTopK > 1 with a tree-capable base.
JetSpec Draft Models#
JetSpec draft checkpoints are detected by jetspec_config in config.json and
exported with DFlashDraftModel using causal proposal attention. The validated
runtime path is branching tree verification: export the base with
--jetspec-tree-base --jetspec-draft-dir <draft_checkpoint>, export the draft
with --jetspec-draft --jetspec-draft-dir <draft_checkpoint>, and run with
--specDraftTopK > 1. --jetspecBlockSize and --dflashBlockSize configure
the same cached-draft proposal block size; the JetSpec spelling is provided for
clarity in JetSpec command lines.
For the listed Qwen3 pair, disable thinking mode in the input JSON when evaluating accuracy, acceptance rate, or throughput.
Draft checkpoint |
Base model |
|---|---|