Overview#
TensorRT Edge-LLM is NVIDIA’s C++ inference runtime for generative models on NVIDIA Jetson, NVIDIA DRIVE, and NVIDIA DGX Spark. It supports text, image, audio, speech, and action workflows while keeping the deployment runtime free of Python dependencies.
See the support matrix for the release software stacks and supported models for checkpoint IDs.
Deployment Workflows#
TensorRT Edge-LLM provides two engine frontends. Both produce artifacts consumed by the same C++ runtimes.
flowchart LR
HF[Hugging Face checkpoint]
Q[Optional quantization]
E[Checkpoint exporter]
O[ONNX components]
C[C++ component builders]
D[Experimental direct builder]
T[TensorRT engines]
R[Model runtime]
HF --> Q
Q --> E --> O --> C --> T
Q --> D --> T
T --> R
Frontend |
Command |
Use it for |
|---|---|---|
ONNX workflow |
|
Supported deployment path, portable intermediate artifacts, and explicit component control |
Direct frontend |
|
Experimental on-device compilation directly from a local checkpoint |
Quantization is optional. Unquantized and supported pre-quantized checkpoints can
be compiled directly; use tensorrt-edgellm-quantize only to create a new
quantized checkpoint.
Runtime Capabilities#
Paged attention, FP8 KV cache, LoRA, streaming, and KV cache reuse
EAGLE3, MTP, DFlash, and DSpark speculative decoding on supported models
Image and audio encoders, speech generation, ASR, and action generation
Model-specific runtimes for pipelines whose I/O contract is not LLM-shaped
Experimental Python API and OpenAI-compatible server over the C++ runtime
Feature availability depends on the model and deployment. In particular, KV cache reuse has a narrower support boundary than ordinary inference.
Start Here#
Select a modality-specific workflow from the examples.
For implementation details, see the checkpoint exporter, direct builder, and C++ runtime design guides.