Announcements#

Release notes, technical updates, examples, and deployment stories from the Model Optimizer team.

September 15, 2026 · Model Optimizer Team

Quantizing a 1.5 TB Kimi-K3 Model on a Single GPU

Layerwise calibration and per-layer shard export drop the memory floor for PTQ from one model to one layer.

quantizationnvfp4layerwisemoesingle-gpumodelopt