Skip to main content
Ctrl+K
Transformer Engine 2.21.0.dev0 - Home Transformer Engine 2.21.0.dev0 - Home

Transformer Engine 2.21.0.dev0

  • GitHub
Transformer Engine 2.21.0.dev0 - Home Transformer Engine 2.21.0.dev0 - Home

Transformer Engine 2.21.0.dev0

  • GitHub

Table of Contents

  • Home

Getting Started

  • Installation
  • Getting Started
  • Datatype and hardware support matrix
  • Frequently Asked Questions (FAQ)
  • Transformer Engine vX.YZ Release Notes

Project

  • Project updates
  • Ecosystem and historical integrations
  • Resources

Python API documentation

  • Common API
  • Framework-specific API
    • PyTorch
    • Jax

Features

  • Low precision training
    • Introduction
    • Performance Considerations
    • FP8 Current Scaling
    • FP8 Delayed Scaling
    • FP8 Blockwise Scaling
    • MXFP8
    • NVFP4
    • Fine-grained quantization recipes
    • GEMM Speedups Across Precisions
  • Other optimizations
    • CPU Offloading

Examples and Tutorials

  • Using FP8 and FP4 with Transformer Engine
  • Performance Optimizations
  • Accelerating Hugging Face Llama 2 and 3 Fine-Tuning with Transformer Engine
  • Accelerating Hugging Face Gemma Inference with Transformer Engine
  • Accelerating Hugging Face Mixtral MoE Fine-Tuning with Transformer Engine
  • Export to ONNX and inference using TensorRT
  • JAX: Integrating TransformerEngine into an existing framework
    • JAX: Dense GEMMs with TransformerEngine
    • JAX: Collective GEMMs with TransformerEngine
    • JAX: Attention with TransformerEngine
      • JAX: Single-GPU Attention with TransformerEngine
      • JAX: Context-Parallel Attention with TransformerEngine
    • JAX: Expert Parallelism with TransformerEngine
  • Operation fuser API
  • GEMM Profiling Tutorial

Advanced

  • C/C++ API
    • transformer_engine.h
    • activation.h
    • cast_transpose_noop.h
    • cast.h
    • cudnn.h
    • fused_attn.h
    • fused_rope.h
    • gemm.h
    • multi_tensor.h
    • normalization.h
    • padding.h
    • permutation.h
    • recipe.h
    • softmax.h
    • swizzle.h
    • transpose.h
  • Precision debug tools
    • Getting started
    • Config File Structure
    • API
      • Setup
      • Debug features
      • Calls to Nvidia-DL-Framework-Inspect
    • Distributed training
    • Adding custom feature to precision debug tools
  • Environment Variables
  • Attention Is All You Need!
  • Deep Dive into CP + THD + AG + Striped>1 + SWA support for Transformer Engine JAX
  • Resources

Resources#

Papers#

  • Attention Is All You Need

  • Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

  • Reducing Activation Recomputation in Large Transformer Models

  • FP8 Formats for Deep Learning

Videos#

  • Stable and Scalable FP8 Deep Learning Training on Blackwell | GTC 2025

  • Blackwell Numerics for AI | GTC 2025

  • Building LLMs: Accelerating Pretraining of Foundational Models with FP8 Precision | GTC 2025

  • From FP8 LLM Training to Inference: Language AI at Scale | GTC 2025

  • What’s New in Transformer Engine and FP8 Training | GTC 2024

  • FP8 Training with Transformer Engine | GTC 2023

  • FP8 for Deep Learning | GTC 2023

  • Inside the Hopper Architecture | GTC 2022

previous

Ecosystem and historical integrations

next

Common API

On this page
  • Papers
  • Videos
NVIDIA NVIDIA
Privacy Policy | Your Privacy Choices | Terms of Service | Accessibility | Corporate Policies | Product Security | Contact

Copyright © 2022-2026, NVIDIA CORPORATION & AFFILIATES. All rights reserved..