Resources# Papers# Attention Is All You Need Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism Reducing Activation Recomputation in Large Transformer Models FP8 Formats for Deep Learning Videos# Stable and Scalable FP8 Deep Learning Training on Blackwell | GTC 2025 Blackwell Numerics for AI | GTC 2025 Building LLMs: Accelerating Pretraining of Foundational Models with FP8 Precision | GTC 2025 From FP8 LLM Training to Inference: Language AI at Scale | GTC 2025 What’s New in Transformer Engine and FP8 Training | GTC 2024 FP8 Training with Transformer Engine | GTC 2023 FP8 for Deep Learning | GTC 2023 Inside the Hopper Architecture | GTC 2022