===============================================
Quantization-Aware Training and Distillation
===============================================
Quantization-aware training (QAT) and quantization-aware distillation (QAD) recover
quality that is lost when a model is quantized. Both train a model with simulated
quantization enabled, so the resulting checkpoint retains its ModelOpt quantization
state and can be exported for deployment.
QAT versus QAD
==============
QAT is standard supervised fine-tuning of a quantized model. The training objective
is the usual cross-entropy (CE) loss against labeled data, while the quantized
forward pass lets the weights adapt to quantization error. Use QAT when the goal is
to adapt a quantized model to a task or dataset.
QAD uses knowledge distillation instead: a frozen BF16 teacher guides the quantized
student with a logit-level KL-divergence loss. It is usually used after PTQ, with
the original BF16 model as the teacher, to recover quality lost specifically to
quantization. QAD requires the teacher model during training and therefore uses more
memory and compute than QAT.
In both cases, begin with a PTQ checkpoint and keep its quantization configuration
unchanged during training. After training, export the quantized checkpoint using the
export workflow for the framework that produced it.
QAD workflow and rationale
==========================
QAD is a two-stage workflow:
#. Start from the BF16 checkpoint and use PTQ to create a quantized student
checkpoint. Because QAD is intended to recover the remaining accuracy gap, this
stage can use a more aggressive recipe than a PTQ-only deployment would accept.
#. Train that PTQ checkpoint against the frozen BF16 teacher. Each student forward
pass uses simulated quantization and the distillation loss aligns the student
logits with the teacher. Export the resulting QAD checkpoint for deployment.
The PTQ recipe determines how quantization scales are handled during QAD. Dynamic
scales from max-calibrated PTQ checkpoints can be recomputed during training;
scales from MSE-based static PTQ checkpoints should remain frozen, because repeating
the scale search each step is prohibitively expensive.
`Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer `_
walks through this PTQ-to-QAD-to-export process, including selection of the student
recipe, training data and sequence length, and scale handling.
The `QAD paper `_ recommends QAD for accuracy
recovery after aggressive quantization, particularly for models that have gone
through multi-stage post-training such as SFT, RL, or model merging. It reports that
QAD is more stable and less complex to engineer than conventional QAT in these
settings, and that the teacher signal makes recovery more robust when training data
quality or coverage is limited. QAT remains the direct choice when task-specific CE
fine-tuning is the objective.
Choose a framework
==================
ModelOpt supports QAT with Hugging Face, Megatron-Bridge, and Megatron-LM.
.. list-table::
:header-rows: 1
:widths: 20 40 40
* - Framework
- Advantages
- Trade-offs
* - Hugging Face
- Starts directly from Hugging Face checkpoints and is the simplest path for
small to medium models. It supports FSDP2, DDP, and DeepSpeed through
Accelerate, with no conversion to Megatron-Core.
- Its parallelism is less efficient for large-scale training, so it is better
suited to smaller models than the Megatron-based options.
* - Megatron-Bridge
- Automatically converts Hugging Face models to Megatron-Core and uses
Megatron-LM's distributed training stack. It is a convenient scalable
workflow without a separate conversion step.
- The high-level workflow exposes fewer customization points than working
directly in Megatron-LM.
* - Megatron-LM
- Provides the most control over model configuration, data, parallelism, and
training behavior, making it the most customizable option for large-model
training.
- Requires manually converting the model to Megatron-Core and managing that
checkpoint workflow.
Run QAT or QAD
==============
Hugging Face
------------
The Hugging Face workflow is manual; there is no launcher example. The
`Hugging Face QAT/QAD examples `_
provide end-to-end QAT and QAD commands. In summary:
#. Quantize the Hugging Face model with ``examples/llm_qat/quantize.py``.
#. For QAT, run ``examples/llm_qat/train.py`` with the quantized checkpoint and a
QAT configuration. This trains with CE loss on the labeled dataset.
#. For QAD, run the same training script with a QAD configuration and
``--teacher_model`` pointing at the BF16 model.
#. Export the trained checkpoint with ``examples/llm_qat/export.py``.
For example, the QAT and QAD training configurations in the examples are
``configs/train/qat_nvfp4.yaml`` and ``configs/train/qad_nvfp4.yaml`` respectively.
The ``Quick Start: QAT (Hugging Face)`` navigation entry links to these examples.
Run the Megatron workflows in a NeMo container
----------------------------------------------
Megatron-Bridge and Megatron-LM require CUDA, Megatron-Core, and their framework
dependencies. The `NeMo container catalog `_
lists the latest available image; the commands below use ``nvcr.io/nvidia/nemo:26.08``.
Run them on a host with NVIDIA GPUs and Docker configured for GPU access.
From the Model Optimizer repository root, prepare a directory for checkpoints and
Hugging Face caches, then start an interactive container. Replace ``$HF_TOKEN`` with
a token that can read the selected model and dataset.
.. code-block:: bash
export MODELOPT_DIR="$PWD"
export QAT_WORK_DIR="$PWD/qat-qad-work"
mkdir -p "$QAT_WORK_DIR/hf-cache"
docker run --rm -it --gpus all --shm-size=16g --net=host --ulimit memlock=-1 \
-e HF_TOKEN="$HF_TOKEN" \
-v "$MODELOPT_DIR":/opt/Model-Optimizer \
-v "$QAT_WORK_DIR":/workspace \
-v "$QAT_WORK_DIR/hf-cache":/root/.cache/huggingface \
-w /opt/Model-Optimizer \
nvcr.io/nvidia/nemo:26.08 bash
Inside the container, authenticate with ``hf auth login --token "$HF_TOKEN"`` if
the token was not provided through the environment. The mounted ``/workspace``
directory preserves checkpoints after the container exits. The examples below use a
small Qwen model and one GPU to demonstrate the command shape; increase the GPU
count and choose TP, PP, CP, and EP to fit the target model and sequence length.
Megatron-Bridge
---------------
Megatron-Bridge converts the Hugging Face model to Megatron-Core automatically.
First create a quantized Megatron checkpoint; it is the input to either QAT or QAD:
.. code-block:: bash
cd /opt/Model-Optimizer/examples/megatron_bridge
export MODEL=Qwen/Qwen3-0.6B
export PTQ_CKPT=/workspace/qwen3-0.6b-nvfp4-megatron
torchrun --nproc_per_node 1 quantize.py \
--hf_model_name_or_path "$MODEL" \
--quant_cfg nvfp4 \
--calib_batch_size 1 \
--calib_num_samples 128 \
--seq_length 4096 \
--export_megatron_path "$PTQ_CKPT"
For QAT, load ``$PTQ_CKPT`` into your Megatron-Bridge SFT configuration, restore
the ModelOpt state, and train with the normal CE objective. This keeps the student
fake-quantized during SFT. The repository's Megatron-Bridge entry point is focused
on distillation; use the framework's SFT application for this CE-only step.
For QAD, run the included ``distill.py`` entry point. ``--student_megatron_path``
restores the quantized student and its ModelOpt state, while the BF16 Hugging Face
model is loaded as the frozen teacher. The following mock-data command verifies the
end-to-end path; replace ``--use_mock_data`` with ``--data_paths`` (or ``--sft`` and
``--sft_dataset_root``) for a real run.
.. code-block:: bash
torchrun --nproc_per_node 1 distill.py \
--teacher_hf_path "$MODEL" \
--student_hf_path "$MODEL" \
--student_megatron_path "$PTQ_CKPT" \
--use_mock_data \
--seq_length 512 \
--mbs 1 \
--gbs 8 \
--train_iters 100 \
--eval_interval 10 \
--eval_iters 4 \
--output_dir /workspace/qwen3-0.6b-nvfp4-qad
Export the final QAD checkpoint to a deployable unified Hugging Face checkpoint:
.. code-block:: bash
torchrun --nproc_per_node 1 export_quantized_megatron_to_hf.py \
--hf_model_name_or_path "$MODEL" \
--megatron_path /workspace/qwen3-0.6b-nvfp4-qad/checkpoints \
--export_unified_hf_path /workspace/qwen3-0.6b-nvfp4-qad-hf
The `Megatron-Bridge README `_
includes the complete PTQ, QAD, data-preparation, and export options.
There is also a complete Slurm launcher example for QAD. It tokenizes the data,
performs NVFP4 PTQ, distills the quantized student, and exports a unified Hugging
Face checkpoint:
.. code-block:: bash
cd tools/launcher
source .env-slurm
uv run launch.py \
--yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/mbridge_qad.yaml \
--yes
See the `Megatron-Bridge launcher configuration `_
before running it, and adjust the model, data paths, output paths, and Slurm topology
for your environment.
Megatron-LM
-----------
For manual QAT, first apply PTQ, then fine-tune the quantized checkpoint with the
ModelOpt-enabled Megatron-LM post-training scripts. QAT uses the standard SFT/CE
objective. From the NeMo container started above, run:
.. code-block:: bash
cd /opt/Model-Optimizer/tools/launcher/modules/Megatron-LM/examples/post_training/modelopt
export MODEL=meta-llama/Llama-3.2-1B-Instruct
export HF_MODEL_CKPT="$MODEL"
export PTQ_CKPT=/workspace/llama-3.2-1b-nvfp4-ptq
TP=1 PP=1 EP=1 ETP=1 \
MLM_MODEL_SAVE="$PTQ_CKPT" \
MLM_EXTRA_ARGS="--calib-size 128" \
./quantize.sh "$MODEL" NVFP4_DEFAULT_CFG
TP=1 PP=1 EP=1 ETP=1 \
MLM_MODEL_CKPT="$PTQ_CKPT" \
MLM_MODEL_SAVE=/workspace/llama-3.2-1b-nvfp4-qat \
DATASET=Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered \
MLM_EXTRA_ARGS="--modelopt-enabled --train-samples 1000 --lr-decay-samples 1000" \
./finetune.sh "$MODEL"
Replace the example ``DATASET`` value with a dataset appropriate for the task. The
second command loads the quantized checkpoint, retains simulated quantization, and
uses the normal supervised objective. Export it after training:
.. code-block:: bash
TP=1 PP=1 EP=1 ETP=1 \
MLM_MODEL_CKPT=/workspace/llama-3.2-1b-nvfp4-qat \
EXPORT_DIR=/workspace/llama-3.2-1b-nvfp4-qat-hf \
./export.sh "$MODEL"
For QAD, manually convert the BF16 Hugging Face model to a Megatron-Core teacher,
then use the PTQ checkpoint as the student. Megatron-LM requires a NeMo-style
``model_config.yaml`` for the teacher checkpoint; place it in the teacher directory
or pass its path with ``--export-kd-teacher-model-config``. The
`Megatron-LM ModelOpt post-training examples `_
are the authoritative reference for the conversion, quantization, and training
configuration.
.. code-block:: bash
export TEACHER_CKPT=/workspace/llama-3.2-1b-bf16-mcore
TP=1 PP=1 EP=1 ETP=1 \
MLM_MODEL_SAVE="$TEACHER_CKPT" \
./convert.sh "$MODEL"
TP=1 PP=1 EP=1 ETP=1 \
MLM_MODEL_CKPT="$PTQ_CKPT" \
MLM_MODEL_SAVE=/workspace/llama-3.2-1b-nvfp4-qad \
DATASET=Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered \
MLM_EXTRA_ARGS="--modelopt-enabled --export-kd-teacher-load $TEACHER_CKPT --train-samples 1000 --lr-decay-samples 1000" \
./finetune.sh "$MODEL"
The second command adds the frozen teacher and the default logit-level distillation
loss to the quantized student training. Use ``./export.sh "$MODEL"`` as in the QAT
flow, replacing ``MLM_MODEL_CKPT`` and ``EXPORT_DIR`` with the QAD output paths.
The following launcher example runs the complete Megatron-LM QAD flow: import the
BF16 model, create an NVFP4 student with PTQ, distill the student, and export it to
Hugging Face format.
.. code-block:: bash
cd tools/launcher
source .env-slurm
uv run launch.py \
--yaml examples/nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16/megatron_lm_qad.yaml \
--yes
Review the `Megatron-LM launcher configuration `_
before launching. In particular, adapt its data, checkpoint paths, and distributed
topology to the target model and cluster.
Additional resources
====================
* `Hugging Face QAT/QAD examples `_
* `Megatron-Bridge examples `_
* `Megatron-LM ModelOpt post-training examples `_
* `Developing Nemotron 3.5 Lightning NVFP4 with QAD Using NVIDIA Model Optimizer `_
* `Quantization-Aware Distillation for NVFP4 Inference Accuracy Recovery `_