Low precision training# Introduction Training in BF16/FP16 Lower precisions Quantizers Performance Considerations Handling transposes Memory usage Fused layers Distributed training FP8 Current Scaling FP8 data type Scaling factors Transpose handling Distributed training Supported devices Examples Quantizer Developer Notes All-gather of columnwise tensors FP8 Delayed Scaling Quantization with delayed scaling factors Amax History Management Distributed Training Supported devices Quantizer FP8 Blockwise Scaling Data Format Handling transposes Distributed training Examples Supported devices Quantizer Developer Notes Swizzle of scaling factors All-gather of columnwise tensors MXFP8 Data Format Handling transposes Distributed training Examples Supported devices Quantizer Developer Notes Swizzling scaling factors All-gather of columnwise tensors NVFP4 Data Format Stochastic Rounding Random Hadamard Transform Handling transposes Distributed training Examples Supported devices Quantizer Developer Notes Swizzling scaling factors All-gather of columnwise tensors Fine-grained quantization recipes Example: mixing MXFP8, NVFP4, and BF16 CustomRecipe and quantizer factory HybridQuantizer Choosing the columnwise source IdentityQuantizer Example: one format per GEMM Validating and optimizing a recipe API reference GEMM Speedups Across Precisions Example: 5B Model on B300 (Blackwell) Example: 5B Model on H200 (Hopper) Speedup Is Shape-Dependent