Installation#

Prerequisites#

  1. Linux x86_64

  2. CUDA 12.1+ (12.8+ for Blackwell support)

  3. NVIDIA Driver supporting CUDA 12.1 or later.

  4. cuDNN 9.12 or later.

If the CUDA Toolkit headers are not available at runtime in a standard installation path, e.g. within CUDA_HOME, set NVTE_CUDA_INCLUDE_PATH in the environment.

Transformer Engine in NGC Containers#

Transformer Engine library is preinstalled in the PyTorch container in versions 22.09 and later on NVIDIA GPU Cloud.

pip - from PyPI#

Transformer Engine can be directly installed from our PyPI, e.g.

pip3 install --no-build-isolation transformer_engine[pytorch]

To obtain the necessary Python bindings for Transformer Engine, the frameworks needed must be explicitly specified as extra dependencies in a comma-separated list (e.g. [jax,pytorch]). Transformer Engine ships wheels for the core library. Source distributions are shipped for the JAX and PyTorch extensions.

The core package from Transformer Engine (without any framework extensions) can be installed via:

pip3 install transformer_engine[core]

By default, this will install the core library compiled for CUDA 12. The cuda major version can be specified by modified the extra dependency to core_cu12 or core_cu13.

pip - from GitHub#

Additional Prerequisites#

  1. [For PyTorch support] PyTorch with GPU support.

  2. [For JAX support] JAX with GPU support, version >= 0.4.7.

Installation (stable release)#

Execute the following command to install the latest stable version of Transformer Engine:

pip3 install --no-build-isolation git+https://github.com/NVIDIA/TransformerEngine.git@stable

This will automatically detect if any supported deep learning frameworks are installed and build Transformer Engine support for them. To explicitly specify frameworks, set the environment variable NVTE_FRAMEWORK to a comma-separated list (e.g. NVTE_FRAMEWORK=jax,pytorch).

Installation (development build)#

Warning

While the development build of Transformer Engine could contain new features not available in the official build yet, it is not supported and so its usage is not recommended for general use.

Execute the following command to install the latest development build of Transformer Engine:

pip3 install --no-build-isolation git+https://github.com/NVIDIA/TransformerEngine.git@main

This will automatically detect if any supported deep learning frameworks are installed and build Transformer Engine support for them. To explicitly specify frameworks, set the environment variable NVTE_FRAMEWORK to a comma-separated list (e.g. NVTE_FRAMEWORK=jax,pytorch). To only build the framework-agnostic C++ API, set NVTE_FRAMEWORK=none.

In order to install a specific PR, execute (after changing NNN to the PR number):

pip3 install --no-build-isolation git+https://github.com/NVIDIA/TransformerEngine.git@refs/pull/NNN/merge

Installation (from source)#

Execute the following commands to install Transformer Engine from source:

# Clone repository, checkout stable branch, clone submodules
git clone --branch stable --recursive https://github.com/NVIDIA/TransformerEngine.git

cd TransformerEngine
export NVTE_FRAMEWORK=pytorch         # Optionally set framework
pip3 install --no-build-isolation .   # Build and install

If the Git repository has already been cloned, make sure to also clone the submodules:

git submodule update --init --recursive

Extra dependencies for testing can be installed by setting the “test” option:

pip3 install --no-build-isolation .[test]

To build the C++ extensions with debug symbols, e.g. with the -g flag:

NVTE_BUILD_DEBUG=1 pip3 install --no-build-isolation .

Troubleshooting#

Common issues and solutions#

  1. ABI compatibility issues

    • Symptoms: ImportError with undefined symbols when importing Transformer Engine.

    • Solution: Ensure PyTorch and Transformer Engine are built with the same C++ ABI setting. Rebuild PyTorch from source with a matching ABI if needed.

    • Context: This is particularly common with pip-installed PyTorch outside of containers.

  2. Missing headers or libraries

    • Symptoms: CMake errors about missing headers such as cudnn.h, cublas_v2.h, or filesystem.

    • Solution: Install the missing development packages or point to the correct locations:

      export CUDA_PATH=/path/to/cuda
      export CUDNN_PATH=/path/to/cudnn
      
    • If CMake cannot find a C++ compiler, set the CXX environment variable.

  3. Build resource issues

    • Symptoms: Compilation hangs, the system freezes, or the build runs out of memory.

    • Solution: Limit parallel builds:

      MAX_JOBS=1 NVTE_BUILD_THREADS_PER_JOB=1 pip install ...
      
  4. Verbose build logging

    Use verbose output to diagnose a build:

    cd transformer_engine
    pip install -v -v -v --no-build-isolation .
    

UV and virtual environments#

  1. Import error

    Ensure the UV environment is active and install with uv pip install --no-build-isolation <package-or-source-directory> instead of installing into the system environment.

  2. cuDNN sublibrary loading failure

    CUDNN_STATUS_SUBLIBRARY_LOADING_FAILED can occur when Transformer Engine is built against the container’s system cuDNN while packages inside the virtual environment install a different nvidia-cudnn-cu12 or nvidia-cudnn-cu13 version. When building from source, point the build and runtime to cuDNN in the virtual environment:

    export CUDNN_PATH=$(pwd)/.venv/lib/python3.12/site-packages/nvidia/cudnn
    export CUDNN_HOME=$CUDNN_PATH
    export LD_LIBRARY_PATH=$CUDNN_PATH/lib:$LD_LIBRARY_PATH
    
  3. Building wheels

    Use uv build --wheel --no-build-isolation -v when building the wheel and uv pip install --no-build-isolation when installing it. Verbose output helps verify that the build is not pulling in a different PyTorch or JAX version from the active environment.

JAX-specific issues#

FFI registration error

If you see No registered implementation for custom call to <some_te_ffi> for platform CUDA, ensure --no-build-isolation is used both when building and installing Transformer Engine.