DRA Driver for NVIDIA GPUs#

Dynamic Resource Allocation (DRA) is a Kubernetes API for flexibly requesting, configuring, and sharing specialized devices such as GPUs. This page describes how to use the GPU Operator to install and manage DRA Driver for NVIDIA GPUs v0.4.1.

Before using the DRA Driver for NVIDIA GPUs, familiarize yourself with the following documentation:

Important

GPU Operator management of DRA is available as a technology preview. Technology preview features are not supported in production environments and are not functionally complete.

Comparison: DRA and Device Plugin#

The DRA Driver for NVIDIA GPUs and the NVIDIA Kubernetes Device Plugin provide alternative mechanisms for allocating NVIDIA GPU resources. The mechanisms do not provide feature parity. A cluster can have either a GPUCluster resource for DRA or a ClusterPolicy resource for the Device Plugin, but not both.

Use DRA for workloads that have the following requirements:

  • Coordinate Multi-Node NVLink workloads by using ComputeDomains.

  • Use attribute-based GPU selection or allocation-specific device configuration instead of node-wide configuration.

  • Configure dynamic NVIDIA Multi-Instance GPU (MIG), CUDA Multi-Process Service (MPS), CUDA time-slicing, or Virtual Function I/O (VFIO) passthrough for individual allocations.

Use the Device Plugin for workloads that have the following requirements:

  • Request nvidia.com/gpu or MIG extended resources from existing Pod specifications and tools.

  • Schedule CUDA time-slicing or MPS replicas as independent extended resources.

  • Use the broader set of components managed through ClusterPolicy, Device Plugin health reporting, or Kubernetes Pod priority and preemption.

Capability Comparison#

The following table compares the GPU allocation capabilities of the two mechanisms when managed by the GPU Operator:

GPU Allocation Capability Comparison#

Capability

DRA with GPUCluster

Device Plugin with ClusterPolicy

API and Device Selection

Uses DRA API objects for structured device selection and per-claim configuration.

Uses Kubernetes extended resource names and counts. Configuration applies at the node or cluster level.

Full GPU and MIG Device Allocation

Allocates full GPU devices and existing MIG devices. Alpha DynamicMIG creates MIG devices for a claim.

Allocates full GPU devices or preconfigured MIG devices. MIG Manager configures MIG geometry before allocation.

GPU Sharing

Alpha TimeSlicingSettings and MPSSupport configure sharing through DRA claims.

Advertises time-slicing replicas and experimental MPS replicas as independent extended resources.

Multi-Node NVLink

Provides the ComputeDomain API and controller for coordinated workloads.

Does not provide the ComputeDomain API.

Device Health Reporting

Alpha NVMLDeviceHealthCheck is disabled by default. gRPC health probes report kubelet plugin availability only.

Reports unhealthy devices through the Kubernetes Device Plugin API.

Device Injection and Managed Components

Requires a Container Device Interface (CDI)-compatible runtime. GPUCluster manages DRA, ComputeDomains, NVIDIA Data Center GPU Manager (DCGM), DCGM Exporter, and the DRA validator.

Supports multiple device injection strategies. ClusterPolicy manages NVIDIA Container Toolkit, GPU Feature Discovery, MIG Manager, the sandbox device plugin, Kata Containers, KubeVirt, and NVIDIA vGPU Manager.

Kubernetes Scheduling

Uses ResourceClaim objects and DRA scheduler integration. Kubernetes does not support preemption for DRA resources.

Uses Kubernetes extended resource scheduling and supports Kubernetes Pod priority and preemption.

DRA Driver Feature Maturity#

GPU Operator support for deploying DRA components and the maturity of capabilities in the upstream DRA driver are separate considerations. Upstream describes ComputeDomains as officially supported rather than generally available (GA). For this comparison, the table treats the terms as equivalent. GPUCluster always enables GPU allocation, can disable ComputeDomains, and passes draDriver.featureGates values to the DRA driver.

The following table groups the DRA driver capabilities by maturity:

DRA Driver Capability Maturity#

Capability

Maturity

Default

Operator Control

ComputeDomains

Officially supported (GA)

Enabled

draDriver.computeDomains.enabled

Full GPU and existing MIG device allocation

Not officially supported

Enabled

Always enabled

ComputeDomainCliques, CrashOnNVLinkFabricErrors, IMEXDaemonsWithDNSNames

Beta

Enabled

draDriver.featureGates

DeviceMetadata, DynamicMIG, MPSSupport, NVMLDeviceHealthCheck, PassthroughSupport, TimeSlicingSettings

Alpha

Disabled

draDriver.featureGates

Refer to the DRA driver v0.4.1 overview and feature-gate definitions for the authoritative support status, maturity, and default values.

DRA Driver Limitations#

Consider the following limitations before selecting DRA driver v0.4.1:

  • ComputeDomains are officially supported, but full GPU and MIG allocation are not yet officially supported upstream.

  • NVMLDeviceHealthCheck is alpha and disabled by default. The gRPC health probe reports kubelet plugin availability, and DCGM provides telemetry. Neither provides DRA allocation health status.

  • This release does not provide scheduler-accounted capacity sharing among independent ResourceClaim objects. Workloads must share one claim instead of requesting independent replicas.

  • Alpha feature gates have compatibility constraints. DynamicMIG conflicts with PassthroughSupport, NVMLDeviceHealthCheck, and MPSSupport. PassthroughSupport conflicts with NVMLDeviceHealthCheck, and DeviceMetadata requires PassthroughSupport.

  • Kubernetes does not support preemption for DRA resources. Injecting DRA-allocated devices requires a preconfigured CDI-compatible container runtime. The managed gpu.nvidia.com DeviceClass does not set spec.extendedResourceName to nvidia.com/gpu.

Overview#

The GPU Operator manages DRA components through the nvidia.com/v1alpha1 GPUCluster custom resource. GPUCluster is a cluster-scoped singleton and its name must be gpu-cluster. The Helm chart creates this resource when you set gpuCluster.deployCR=true.

A cluster can have either a GPUCluster or a ClusterPolicy resource, but not both. The GPU Operator uses ClusterPolicy to manage components for Device Plugin-based allocation and GPUCluster to manage components for DRA-based allocation.

For GPUCluster, the GPU Operator manages the following components:

  • The DRA driver GPU capability for gpu.nvidia.com, mig.nvidia.com, and vfio.gpu.nvidia.com devices. This capability is always enabled.

  • The ComputeDomain controller and kubelet plugin for Multi-Node NVLink (MNNVL). ComputeDomain support is enabled by default and can be disabled.

  • A DRA validator that allocates a GPU by using a ResourceClaim and verifies that the device is usable.

  • DCGM Exporter, which is enabled by default.

  • Standalone DCGM, which is disabled by default.

The GPUCluster resource does not manage the NVIDIA GPU driver. Use an NVIDIADriver resource to install a containerized driver or use a driver that is pre-installed on the host. The Operator automatically assigns the nvidia.com/gpu.deploy.* labels that control operand placement.

GPUCluster Limitations#

  • Use this workflow for a new installation. An in-place migration from a standalone DRA driver Helm release or from ClusterPolicy to GPUCluster is not supported.

  • GPUCluster does not expose the controller affinity, priority class, toleration, and kubelet-plugin node selector overrides that were used by the previous standalone DRA driver procedure for Google Kubernetes Engine. This page does not provide a managed DRA installation procedure for GKE.

Prerequisites#

In addition to ensuring that your GPUs and cluster align with the GPU Operator support matrix, verify the following prerequisites:

  • Kubernetes v1.34.2 or later with a resource.k8s.io DeviceClass API available.

  • NVIDIA GPU driver version 580 or later.

  • An underlying container runtime that supports CDI and is configured to use CDI for injecting DRA-allocated devices.

  • No ClusterPolicy resource exists in the cluster.

    $ kubectl get clusterpolicy
    

    If a ClusterPolicy exists, use a new cluster for the GPUCluster workflow.

Note

To use an extended-resource request with the DRA driver, enable the DRAExtendedResource feature gate. This feature gate is enabled by default in Kubernetes v1.36.0 and later. Request deviceclass.resource.kubernetes.io/gpu.nvidia.com to use the managed gpu.nvidia.com DeviceClass. The legacy name nvidia.com/gpu requires a separate DeviceClass that sets spec.extendedResourceName to that value.

For ComputeDomain support, verify the following additional prerequisites:

  • NVIDIA Grace Blackwell GPUs with Multi-Node NVLink available on your cluster, such as NVIDIA HGX GB200 NVL72 or NVIDIA HGX GB300 NVL72. Refer to the NVIDIA Multi-Node NVLink Systems documentation for more information.

  • If you use a pre-installed GPU driver, install the corresponding nvidia-imex-* packages through the Linux distribution package manager.

  • If you use a pre-installed GPU driver, disable and mask the IMEX systemd service on every GPU node before installing the GPU Operator:

    $ systemctl disable --now nvidia-imex.service
    $ systemctl mask nvidia-imex.service
    

Install#

The following procedures install the GPU Operator and create the gpu-cluster resource. The Operator deploys the DRA driver operands in the GPU Operator namespace.

  1. Add the NVIDIA Helm repository:

    $ helm repo add nvidia https://helm.ngc.nvidia.com/nvidia \
        && helm repo update
    
  2. Install the GPU Operator by using the procedure for your driver configuration.

    $ helm upgrade --install gpu-operator nvidia/gpu-operator \
        --version=v26.3.3 \
        --namespace gpu-operator \
        --create-namespace \
        --set clusterPolicy.deployCR=false \
        --set gpuCluster.deployCR=true \
        --set driver.nvidiaDriverCRD.enabled=true
    

    The chart creates the default NVIDIADriver resource to manage the NVIDIA GPU driver.

    $ helm upgrade --install gpu-operator nvidia/gpu-operator \
        --version=v26.3.3 \
        --namespace gpu-operator \
        --create-namespace \
        --set clusterPolicy.deployCR=false \
        --set gpuCluster.deployCR=true \
        --set driver.enabled=false
    

By default, the Operator enables both GPU allocation and ComputeDomain support. To install GPU allocation without ComputeDomain support, add the following option to the Helm command:

--set draDriver.computeDomains.enabled=false

Do not install a separate dra-driver-nvidia-gpu Helm release for this managed workflow.

Configure DRA Components#

The GPU Operator Helm values render the specification of the gpu-cluster resource. The following settings provide the primary configuration surface:

Helm value

Description

draDriver.repository, draDriver.image, and draDriver.version

Configure the DRA driver container image.

draDriver.imagePullPolicy and draDriver.imagePullSecrets

Configure image pulling for the DRA driver containers.

draDriver.featureGates

Enable or disable DRA driver feature gates. The Operator renders the map as the FEATURE_GATES environment variable for DRA driver containers.

draDriver.gpus.kubeletPlugin

Configure environment variables, resource requests and limits, and the gRPC health check for the GPU kubelet-plugin container.

draDriver.computeDomains.enabled

Enable or disable the ComputeDomain controller and kubelet-plugin container.

draDriver.computeDomains.controller

Configure environment variables and resource requests and limits for the ComputeDomain controller.

draDriver.computeDomains.kubeletPlugin

Configure environment variables, resource requests and limits, and the gRPC health check for the ComputeDomain kubelet-plugin container.

hostPaths.kubeletRootDir

Configure the kubelet root directory when it differs from /var/lib/kubelet.

daemonsets

Configure common labels, annotations, and tolerations. daemonsets.priorityClassName applies to the DRA driver kubelet plugin and ComputeDomain controller.

For example, the following values enable a DRA driver feature gate, configure container resources, and enable the default gRPC health checks on their default ports:

draDriver:
  featureGates:
    NVMLDeviceHealthCheck: true
  gpus:
    kubeletPlugin:
      resources:
        requests:
          cpu: 50m
          memory: 64Mi
      healthcheck:
        enabled: true
  computeDomains:
    enabled: true
    controller:
      resources:
        requests:
          cpu: 50m
          memory: 64Mi
    kubeletPlugin:
      healthcheck:
        enabled: true

The GPU kubelet-plugin health-check port defaults to 51516 and the ComputeDomain kubelet-plugin health-check port defaults to 51515. The Operator also enables ComputeDomain GPU clique labeling by setting GPU_CLIQUE_LABEL_ENABLED=true automatically.

Validate Installation#

Reconciliation typically completes within 3 minutes. During reconciliation, the STATUS column progresses from empty to notReady to ready.

  1. Verify that the GPUCluster resource is ready:

    $ kubectl get gpucluster gpu-cluster
    

    Example Output

    NAME          STATUS   AGE
    gpu-cluster   ready    3m
    

    If the resource does not become ready, inspect its conditions and recent events:

    $ kubectl describe gpucluster gpu-cluster
    

    A PrerequisiteNotMet condition can indicate that a ClusterPolicy resource exists. An OperandNotReady condition indicates that the Operator is waiting for one or more managed pods.

  2. Confirm that the managed components are running in the GPU Operator namespace:

    $ kubectl get pods -n gpu-operator
    

    Expected workload names include the following:

    • nvidia-dra-driver-kubelet-plugin

    • nvidia-dra-driver-controller when ComputeDomain support is enabled

    • nvidia-dra-validator

    • nvidia-dcgm-exporter-dra when DCGM Exporter is enabled

    • nvidia-dcgm-dra when standalone DCGM is enabled

    The DRA kubelet-plugin DaemonSet runs a GPU container and, when ComputeDomain support is enabled, a ComputeDomain container. The DRA validator and enabled telemetry operands run one pod on each GPU node.

  3. Verify that the DeviceClasses are available:

    $ kubectl get deviceclass
    

    Example Output

    NAME                                        AGE
    compute-domain-daemon.nvidia.com            2m
    compute-domain-default-channel.nvidia.com   2m
    gpu.nvidia.com                              2m
    mig.nvidia.com                              2m
    vfio.gpu.nvidia.com                         2m
    

    The ComputeDomain DeviceClasses are present only when ComputeDomain support is enabled.

Additional validation procedures are available in the upstream DRA Driver documentation:

Telemetry#

DCGM Exporter is enabled by default with GPUCluster and uses its embedded host engine. Set dcgm.enabled=true to deploy the standalone nvidia-dcgm-dra host engine instead.

When either dcgmExporter.enablePodLabels or dcgmExporter.enablePodUID is enabled, the Operator enables DRA ResourceSlice attribution in DCGM Exporter and grants the exporter read access to ResourceSlice objects. This enables pod metadata enrichment for GPUs allocated through DRA.

Upgrade#

The DRA driver is upgraded as part of the GPU Operator release. When you upgrade the GPU Operator, preserve clusterPolicy.deployCR=false and gpuCluster.deployCR=true in your values. You can select a different DRA driver version by setting draDriver.version. Refer to Upgrading the NVIDIA GPU Operator for the Operator upgrade procedure and required CRD updates.

During an Operator-managed NVIDIA GPU driver upgrade, pods with allocated gpu.nvidia.com ResourceClaims are treated as GPU workloads. The driver upgrade policy applies to those pods, and the DRA kubelet plugin remains available while claims are unprepared. You do not need to configure a separate node label for DRA workload eviction.

This procedure does not migrate a standalone DRA driver Helm release to GPUCluster. For a standalone installation, refer to the upstream DRA driver upgrade guide or select the GPU Operator documentation version that matches the installed release.

Uninstall#

Before uninstalling the GPU Operator, delete user workloads and user-created ResourceClaims for GPUs and verify that the workload pods terminate and the claims are unprepared. Afterward, uninstall the GPU Operator by using Helm.

The chart includes a pre-delete hook that deletes gpu-cluster and waits for the GPUCluster finalizer to perform ordered teardown.

Tip

The uninstall command output can report that the gpu-cluster resource was kept.

However, a pre-delete hook actually deletes the resource.

These resources were kept due to the resource policy:
[GPUCluster] gpu-cluster

release "gpu-operator" uninstalled

The finalizer removes Operator-managed DaemonSets that consume ResourceClaims before the DRA kubelet plugin is removed. Do not use helm uninstall --no-hooks while gpu-cluster exists because Helm can remove the Operator before this teardown completes.

Refer to Uninstalling the GPU Operator for the complete uninstall procedure and CRD cleanup information.

Troubleshooting#

Recover From an Uninstall That Skipped the Pre-Delete Hook#

If you run helm uninstall gpu-operator --no-hooks while gpu-cluster exists, Helm removes the Operator before the pre-delete hook can perform an ordered teardown. The following resources are left in the cluster with no controller to reconcile them:

  • The gpu-cluster resource, which cannot be deleted because the gpucluster.nvidia.com/dra-resourceclaim finalizer requires the Operator.

  • The DRA driver kubelet plugin, DRA validator, and DCGM Exporter pods.

  • Operator-managed ResourceClaims for the preceding pods.

To recover, remove the finalizer from gpu-cluster, then force-delete the orphaned operands:

  1. Remove the finalizer from gpu-cluster:

    $ kubectl patch gpucluster gpu-cluster --type=json \
        -p='[{"op":"remove","path":"/metadata/finalizers"}]'
    
  2. Remove finalizers from the orphaned ResourceClaims in the GPU Operator namespace:

    $ for rc in $(kubectl get resourceclaim -n gpu-operator -o name); do
        kubectl patch $rc -n gpu-operator --type=json \
            -p='[{"op":"remove","path":"/metadata/finalizers"}]'
      done
    
  3. Force-delete any pods that remain in the Terminating state:

    $ kubectl delete pods -n gpu-operator --all --grace-period=0 --force
    
  4. Delete the GPU Operator namespace to remove residual state:

    $ kubectl delete namespace gpu-operator
    

To avoid this recovery path, run helm uninstall without --no-hooks so that the pre-delete hook can complete ordered teardown.

Clean Up ResourceClaims Stuck in deleted,allocated,reserved#

ResourceClaims in the deleted,allocated,reserved state indicate that the API server accepted a delete request, but the DRA driver kubelet plugin has not completed the device _unprepare_ step for the claim. This state can persist if the DRA driver kubelet plugin was removed or restarted before it could unprepare the allocated devices.

Take one of the following actions:

  • If the DRA driver kubelet plugin can be restarted, ensure that the nvidia-dra-driver-kubelet-plugin pod is running. The plugin completes the unprepare step and the ResourceClaims are removed.

  • If the DRA driver kubelet plugin cannot be restarted, such as after an uninstall that skipped the pre-delete hook, remove the finalizers from the affected ResourceClaims:

    $ for rc in $(kubectl get resourceclaim -n gpu-operator -o name); do
        kubectl patch $rc -n gpu-operator --type=json \
            -p='[{"op":"remove","path":"/metadata/finalizers"}]'
      done
    

    Removing finalizers bypasses the unprepare step. Use this option only during recovery when no DRA workloads are running.

Additional Documentation#

For more details on the DRA Driver for NVIDIA GPUs, refer to the following resources: