Low precision training# Introduction Training in BF16/FP16 Lower precisions Performance Considerations Handling transposes Memory usage Fused layers Distributed training FP8 Current Scaling FP8 data type Scaling factors Transpose handling Distributed training Supported devices Examples Developer Notes All-gather of columnwise tensors FP8 Delayed Scaling Quantization with delayed scaling factors Amax History Management Distributed Training Supported devices FP8 Blockwise Scaling Data Format Handling transposes Distributed training Examples Supported devices Developer Notes Swizzle of scaling factors All-gather of columnwise tensors MXFP8 Data Format Handling transposes Distributed training Examples Supported devices Developer Notes Swizzling scaling factors All-gather of columnwise tensors NVFP4 Data Format Stochastic Rounding Random Hadamard Transform Handling transposes Distributed training Examples Supported devices Developer Notes Swizzling scaling factors All-gather of columnwise tensors GEMM Speedups Across Precisions Example: 5B Model on B300 (Blackwell) Example: 5B Model on H200 (Hopper) Speedup Is Shape-Dependent