Unified HuggingFace Checkpoint

We support exporting modelopt-optimized Hugging Face models (transformers and diffusers pipelines/components) and Megatron Core models to a unified checkpoint format that can be deployed in various inference frameworks such as TensorRT-LLM, vLLM, and SGLang.

The workflow is as follows:

  1. Load the Huggingface models or Megatron Core models, quantize with modelopt , and export to the unified checkpoint format, where the layer structures and tensor names are aligned with the original checkpoint.

  2. Load the unified checkpoint in the supported inference framework for accelerated inference.

Export Quantized Model

The modelopt quantized model can be exported to the unified checkpoint format stored as

  1. A group of safetensors files, containing quantized model weights and scaling factors.

  2. A hf_quant_config.json file containing quantization configurations.

  3. Other json files that store the model structure information, tokenizer information, and metadata.

The export API (export_hf_checkpoint) can be used as follows:

from modelopt.torch.export import export_hf_checkpoint

with torch.inference_mode():
    export_hf_checkpoint(
        model,  # The quantized model.
        export_dir,  # The directory where the exported files will be stored.
    )

Note

export_hf_checkpoint also supports diffusers pipelines and components (e.g., UNet/transformer). See the diffusers quantization examples for end-to-end workflows and CLI usage.

Deployment Support Matrix

Supported Quantization Formats

The unified HF export API supports the following quantization formats:

  1. FP8 - 8-bit floating point

  2. FP8_PB - 8-bit floating point with per-block scaling

  3. NVFP4 - NVIDIA 4-bit floating point

  4. NVFP4_AWQ - NVIDIA 4-bit floating point with AWQ optimization

  5. INT4_AWQ - 4-bit integer with AWQ optimization

  6. W4A8_AWQ - 4-bit weights and 8-bit activations with AWQ optimization

Minimum Framework Versions

Framework

Minimum version

TensorRT-LLM

v1.2.0

vLLM

v0.10.1

SGLang

v0.4.10

These are the oldest versions expected to load a unified HF checkpoint. The matrix below is validated against the versions CI currently runs, which are newer — see the container tags in .github/workflows/.

Model Support Matrix

Legend:

  • ✅ — covered by the release deployment test suite (tests/examples/hf_ptq/test_deploy.py), which loads the exported checkpoint in the framework and runs generation.

  • ⚠ — documented as working previously but not in the current test suite; expected to work, unvalidated.

  • - — not currently covered. It may still work; see Models not listed here.

Language models

Model

Quant format

TensorRT-LLM

vLLM

SGLang

Llama 3.1, 3.3

FP8, NVFP4

Llama 4 Scout, Maverick

FP8

Llama 4 Scout

NVFP4

Llama Nemotron Super 49B v1, v1.5

FP8

Llama Nemotron Ultra 253B v1

FP8

Nemotron 3 Nano 30B-A3B

FP8, NVFP4

Nemotron 3 Super 120B-A12B

FP8, NVFP4

Nemotron 3 Ultra 550B-A55B

NVFP4

DeepSeek R1, R1-0528

NVFP4

DeepSeek V3, V3.1, V3.2

NVFP4

DeepSeek V4 Flash

NVFP4

DeepSeek V4 Pro

NVFP4

-

Qwen 3 (8B, 14B, 32B)

FP8, NVFP4

Qwen 3 MoE 235B-A22B

FP8, NVFP4

Qwen 3 MoE 30B-A3B

NVFP4

Qwen 3 Coder 480B-A35B

NVFP4

Qwen 3-Next 80B-A3B

NVFP4

Qwen 3.5 397B-A17B

NVFP4

Qwen 3.5 122B-A10B, Qwen 3.6 35B-A3B

NVFP4

-

-

Qwen 2.5

FP8

Qwen 2.5

NVFP4

-

QwQ-32B

FP8

QwQ-32B

NVFP4

-

Phi-4 reasoning-plus

FP8, NVFP4

Gemma 4 31B

NVFP4

Gemma 4 26B-A4B

NVFP4

-

-

GLM-4.7, GLM-5, GLM-5.2

NVFP4

GLM-5.1

NVFP4

-

Kimi K2-Thinking, K2.5

NVFP4

Kimi K2.6

NVFP4

-

-

MiniMax M2.5, M3

NVFP4

Mixtral 8x7B

FP8

Mixtral 8x7B

NVFP4

-

-

Vision-language and multimodal models

For VLMs, modelopt quantizes the language model only; the vision encoder is kept in high precision. The exported checkpoint therefore relies on the serving framework’s own multimodal support for that architecture — see the TensorRT-LLM multimodal support matrix.

Model

Quant format

TensorRT-LLM

vLLM

SGLang

Qwen 2.5-VL 7B

FP8, NVFP4

Qwen 3-VL 235B-A22B

NVFP4

Phi-4-multimodal

FP8, NVFP4

Nemotron 3 Nano Omni 30B-A3B

FP8, NVFP4

Speculative decoding drafters

Drafters are deployed on top of their base checkpoint. vLLM is not currently covered for these.

Unlike the tables above, drafter coverage is conditional: the EAGLE3 tests run only when MODELOPT_LOCAL_EAGLE_MODEL points at a directory containing the drafter, and skip otherwise. ✅ below means “covered when that drafter is available locally”.

Drafter

Quant format

TensorRT-LLM

vLLM

SGLang

EAGLE3 for Llama 3.3 70B, Llama 4 Maverick

FP8

-

EAGLE3 for Qwen 3 235B-A22B (incl. Thinking-2507, FP4)

BF16, NVFP4

-

EAGLE3 for Qwen 3 30B-A3B-Thinking-2507

BF16

-

EAGLE3 for Kimi K2-Thinking, K2.5, K2.6

NVFP4

-

EAGLE3 for gpt-oss-120b

BF16

-

Medusa for Llama 3.1 8B

FP8

-

Diffusion models

Model

Quant format

TensorRT-LLM

vLLM

SGLang

Wan 2.2 T2V A14B

FP8, NVFP4

-

DiffusionGemma 26B-A4B

NVFP4

Note

NVFP4 inference requires Blackwell GPUs. Hopper can produce an NVFP4 checkpoint but cannot serve it. On B300/GB300 (sm_103) use a CUDA-13 build of the serving framework; CUDA-12 builds lack the sm_103 FP4 kernels.

Models not listed here

This matrix records the combinations modelopt validates. It is not an exhaustive list of what will run: vLLM, SGLang, and TensorRT-LLM load unified HF checkpoints generically, so a model built from standard nn.Linear layers with an hf_quant_config.json will often deploy without any modelopt change. Check the serving framework’s own model support list first, then try it.

The exact checkpoints behind every Y above, including tensor-parallel size and minimum SM version, are listed in tests/examples/hf_ptq/test_deploy.py; most are published under the NVIDIA Hugging Face organization.

Deployment with Selected Inference Frameworks

Follow the TensorRT-LLM installation instructions.

FP8 and NVFP4 quantized models are supported; you need v1.2.0 or later version of TensorRT-LLM.

To run modelopt quantized model from Huggingface model hub, e.g., nvidia/Llama-3.1-8B-Instruct-FP8, refer to the sample code below:

from tensorrt_llm import LLM, SamplingParams

def main():

    prompts = [
        "Hello, my name is",
        "The president of the United States is",
        "The capital of France is",
        "The future of AI is",
    ]
    sampling_params = SamplingParams(temperature=0.8, top_p=0.95)

    llm = LLM(model="nvidia/Llama-3.1-8B-Instruct-FP8")

    outputs = llm.generate(prompts, sampling_params)

    for output in outputs:
        prompt = output.prompt
        generated_text = output.outputs[0].text
        print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")

if __name__ == '__main__':
    main()

Follow vLLM installation instructions.

FP8 and NVFP4 quantized models are supported; you need v0.10.1 or later version of vLLM. Pass quantization="modelopt" for FP8 and quantization="modelopt_fp4" for NVFP4.

To run modelopt quantized model from Huggingface model hub, e.g., nvidia/Llama-3.1-8B-Instruct-FP8, refer to the sample code below:

from vllm import LLM, SamplingParams

def main():

    model_id = "nvidia/Llama-3.1-8B-Instruct-FP8"
    sampling_params = SamplingParams(temperature=0.8, top_p=0.9)

    prompts = [
        "Hello, my name is",
        "The president of the United States is",
        "The capital of France is",
        "The future of AI is",
    ]

    llm = LLM(model=model_id, quantization="modelopt")
    outputs = llm.generate(prompts, sampling_params)

    for output in outputs:
        prompt = output.prompt
        generated_text = output.outputs[0].text
        print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")

if __name__ == "__main__":
    main()

Follow the SGLang installation instructions.

FP8 and NVFP4 quantized models are supported; you need v0.4.10 or later version of SGLang. Pass quantization="modelopt" for FP8 and quantization="modelopt_fp4" for NVFP4.

To run modelopt quantized model from Huggingface model hub, e.g., nvidia/Llama-3.1-8B-Instruct-FP8, refer to the sample code below:

import sglang as sgl

def main():

    prompts = [
        "Hello, my name is",
        "The president of the United States is",
        "The capital of France is",
        "The future of AI is",
    ]
    sampling_params = {"temperature": 0.8, "top_p": 0.95}
    llm = sgl.Engine(model_path="nvidia/Llama-3.1-8B-Instruct-FP8", quantization="modelopt")

    outputs = llm.generate(prompts, sampling_params)
    for prompt, output in zip(prompts, outputs):
        print("===============================")
        print(f"Prompt: {prompt}\nGenerated text: {output['text']}")

if __name__ == "__main__":
    main()