Announcements#
Release notes, technical updates, examples, and deployment stories from the Model Optimizer team.
Local Hessian: Better NVFP4 Weight Scales from Layer Inputs
How Local Hessian sets NVFP4 block scales from layer inputs, and how it measures up against max, MSE, Four-over-six and GPTQ.
AutoQuantize: A Fast Automatic Mixed-Precision Assignment
AutoQuantize finds low-sensitivity mixed-precision assignments with gradient-based scoring under a modeled effective-bits budget.
Model Optimizer announcements are moving to GitHub Pages
The GitHub Pages site now starts with announcements while the existing API documentation remains available in the docs navigation.
DSpark vs Domino: Same DFlash Backbone, Different Correction Heads
DSpark and Domino share a DFlash backbone but make different correction-head tradeoffs: ModelOpt's vanilla Markov head versus a GRU.
No announcements match this search.