Supported Models#
Supported checkpoint IDs are listed below. See the support matrix for platform coverage and the examples for workflows.
Text Generation#
Llama 3.x
Original:
Pre-quantized:
Qwen2 / Qwen2.5
Original:
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B, deepseek-ai/DeepSeek-R1-Distill-Qwen-7B, deepseek-ai/DeepSeek-R1-Distill-Qwen-14B
Pre-quantized:
Qwen/Qwen2-0.5B-Instruct-AWQ, Qwen/Qwen2-0.5B-Instruct-GPTQ-Int4
Qwen/Qwen2-1.5B-Instruct-AWQ, Qwen/Qwen2-1.5B-Instruct-GPTQ-Int4
Qwen/Qwen2-7B-Instruct-AWQ, Qwen/Qwen2-7B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-0.5B-Instruct-AWQ, Qwen/Qwen2.5-0.5B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-1.5B-Instruct-AWQ, Qwen/Qwen2.5-1.5B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-3B-Instruct-AWQ, Qwen/Qwen2.5-3B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-7B-Instruct-AWQ, Qwen/Qwen2.5-7B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-14B-Instruct-AWQ, Qwen/Qwen2.5-14B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-0.5B-Instruct-AWQ, Qwen/Qwen2.5-Coder-0.5B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-1.5B-Instruct-AWQ, Qwen/Qwen2.5-Coder-1.5B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-3B-Instruct-AWQ, Qwen/Qwen2.5-Coder-3B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-7B-Instruct-AWQ, Qwen/Qwen2.5-Coder-7B-Instruct-GPTQ-Int4
Qwen/Qwen2.5-Coder-14B-Instruct-AWQ, Qwen/Qwen2.5-Coder-14B-Instruct-GPTQ-Int4
Qwen3
Original:
Pre-quantized:
Qwen vision-language families
Qwen2.5-VL:
Qwen/Qwen2.5-VL-3B-Instruct-AWQ, Qwen/Qwen2.5-VL-7B-Instruct-AWQ
nvidia/Qwen2.5-VL-7B-Instruct-FP8, nvidia/Qwen2.5-VL-7B-Instruct-NVFP4
Qwen3-VL and compatible checkpoints:
Qwen3.5 / Qwen3.6
InternVL and Phi-4
InternVL:
OpenGVLab/InternVL3-8B-hf, OpenGVLab/InternVL3-9B, OpenGVLab/InternVL3-9B-Instruct
OpenGVLab/InternVL3_5-1B-HF, OpenGVLab/InternVL3_5-2B-HF, OpenGVLab/InternVL3_5-4B-HF
OpenGVLab/InternVL3-1B-AWQ, OpenGVLab/InternVL3-2B-AWQ, OpenGVLab/InternVL3-8B-AWQ, OpenGVLab/InternVL3-14B-AWQ
Phi-4:
Nemotron-3 / 3.5
Gemma4
Text, image, and audio input:
Text and image input:
MTP assistants:
DiffusionGemma
Speech Recognition#
Qwen3-ASR: Qwen/Qwen3-ASR-0.6B, Qwen/Qwen3-ASR-1.7B
Nemotron-3.5-ASR: nvidia/nemotron-3.5-asr-streaming-0.6b
Speech Generation#
Action and Multimodal Reasoning#
Alpamayo: nvidia/Alpamayo-R1-10B
Cosmos3-Edge: nvidia/Cosmos3-Edge, nvidia/Cosmos3-Edge-Policy-DROID
Speculative Draft Checkpoints#
EAGLE3 Draft Models#
Draft checkpoint |
Base model |
|---|---|
DFlash Draft Models#
Draft checkpoint |
Base model |
|---|---|