Announcements#
Release notes, technical updates, examples, and deployment stories from the Model Optimizer team.
Recovering W4A4 NVFP4 Accuracy with Quantization-Aware Distillation
Weight-only NVFP4 is slower than BF16 on Blackwell. W4A4 beats it in 9 of 12 shapes, and QAD recovers the accuracy W4A4 costs on Qwen3.6-35B-A3B.
Improving NVFP4 Accuracy with Local-Hessian Weight Scales
How Local-Hessian selects NVFP4 block scales to minimize layer output error, and how it compares with max, MSE, Four-over-six, and GPTQ.
AutoQuantize: A Fast Automatic Mixed-Precision Assignment
AutoQuantize finds low-sensitivity mixed-precision assignments with gradient-based scoring under a modeled effective-bits budget.
Model Optimizer announcements are moving to GitHub Pages
The GitHub Pages site now starts with announcements while the existing API documentation remains available in the docs navigation.
DSpark vs Domino: Same DFlash Backbone, Different Correction Heads
DSpark and Domino share a DFlash backbone but make different correction-head tradeoffs: ModelOpt's vanilla Markov head versus a GRU.
No announcements match this search.