Announcements#

Release notes, technical updates, examples, and deployment stories from the Model Optimizer team.

September 28, 2026 · Model Optimizer Team

Quantizing a 4.9 TB Qwen3.8 Model on a Single GPU

Layerwise calibration and per-layer shard export drop the memory floor for PTQ from one model to one layer: a 4.9 TB Qwen3.8 checkpoint quantized on a single GB300.

quantizationnvfp4layerwisemoesingle-gpumodelopt
September 9, 2026 · Model Optimizer Team

Improving NVFP4 Accuracy with Local-Hessian Weight Scales

How Local-Hessian selects NVFP4 block scales to minimize layer output error, and how it compares with max, MSE, Four-over-six, and GPTQ.

local-hessianquantizationnvfp4calibrationmodelopt