Project updates#
This page archives Transformer Engine releases, technical articles, research, and community news. The README highlights only the most recent updates.
2026#
[09/2026] Transformer Engine v2.19 adds Rubin support, hybrid quantization, MXFP8 EP communication, and expanded FP8 attention support.
[09/2026] Accelerating Dropless MoE Training in JAX with NVIDIA Transformer Engine.
[08/2026] Transformer Engine v2.18.
[07/2026] Transformer Engine v2.17.
[06/2026] Boosting MoE Training Throughput with Advanced Fusion Kernels.
[06/2026] Train Models Faster with JAX and MaxText Using NVFP4 on NVIDIA Blackwell.
[04/2026] Run High-Throughput Reinforcement Learning Training with End-to-End FP8 Precision.
[02/2026] Using NVFP4 Low-Precision Model Training for Higher Throughput Without Losing Accuracy.
2025#
[12/2025] NVIDIA Nemotron 3: Efficient and Open Intelligence — trained with NVFP4 on Transformer Engine.
[11/2025] NVIDIA Blackwell Architecture Sweeps MLPerf Training v5.1 Benchmarks.
[11/2025] Scale Biology Transformer Models with PyTorch and NVIDIA BioNeMo Recipes.
[11/2025] FP8 Training of Large-Scale RL Models.
[09/2025] Pretraining Large Language Models with NVFP4.
[09/2025] Native FP8 Mixed Precision Training for Ling 2.0, Open Sourced!.
[09/2025] Faster Training Throughput in FP8 Precision with NVIDIA NeMo.
[08/2025] How We Built DeepL’s Next-Generation LLMs with FP8 for Training and Inference.
[08/2025] NVFP4 Trains with Precision of 16-Bit and Speed and Efficiency of 4-Bit.
[06/2025] Floating Point 8: An Introduction to Efficient, Lower-Precision AI Training.
[05/2025] Advanced Optimization Strategies for LLM Training on NVIDIA Grace Hopper.
[03/2025] Stable and Scalable FP8 Deep Learning Training on Blackwell.
[03/2025] Measure and Improve AI Workload Performance with NVIDIA DGX Cloud Benchmarking.
[02/2025] Understanding the Language of Life’s Biomolecules Across Evolution at a New Scale with Evo 2.
[02/2025] NVIDIA DGX Cloud Introduces Ready-To-Use Templates to Benchmark AI Platform Performance.
2024 and earlier#
[11/2024] Developing a 172B LLM with Strong Japanese Capabilities Using NVIDIA Megatron-LM.
[11/2024] How FP8 Boosts LLM Training by 18% on Amazon SageMaker P5 Instances.
[11/2024] Efficiently Train Models with Large Sequence Lengths Using Amazon SageMaker Model Parallel.
[03/2024] Turbocharged Training: Optimizing the Databricks Mosaic AI Stack with FP8.
[03/2024] FP8 Training Support in SageMaker Model Parallelism Library.
[12/2023] New NVIDIA NeMo Framework Features and NVIDIA H200.
[11/2023] Inflection-2: The Next Step Up.
[11/2023] Unleashing the Power of Transformers with NVIDIA Transformer Engine.
[09/2023] Transformer Engine Added to the AWS Deep Learning Container for PyTorch Training.
[06/2023] Breaking MLPerf Training Records with NVIDIA H100 GPUs.
[04/2023] Benchmarking Large Language Models on NVIDIA H100 GPUs with CoreWeave.