Installing the NVIDIA GPU Operator with DRA on OpenShift#
Introduction#
Dynamic Resource Allocation (DRA) is a Kubernetes API for requesting, configuring, and sharing specialized devices such as GPUs.
Starting with NVIDIA GPU Operator v26.7, you can use the GPUCluster and NVIDIADriver custom resources to enable DRA-based GPU allocation on Red Hat OpenShift Container Platform, as an alternative to the Device Plugin-based ClusterPolicy workflow described in Installing the NVIDIA GPU Operator on OpenShift.
Important
A cluster can have either a GPUCluster resource for DRA or a ClusterPolicy resource for the Device Plugin, but not both.
This page assumes a new installation. Refer to DRA Driver for NVIDIA GPUs for the full DRA vs. Device Plugin capability comparison, feature-maturity matrix, and limitations, which apply equally to OpenShift.
Note
This procedure requires GPU Operator v26.7 or later and Kubernetes v1.34.2 or later, which corresponds to Red Hat OpenShift Container Platform 4.21 or later. This guide is validated on Red Hat OpenShift Container Platform 4.22.
Prerequisites#
A working OpenShift cluster with a GPU worker node. Refer to Prerequisites for GPU Operator on OpenShift.
The Node Feature Discovery (NFD) Operator installed and running. Refer to Installing the Node Feature Discovery Operator on OpenShift.
No
ClusterPolicyresource exists in the cluster:$ oc get clusterpolicyIf a
ClusterPolicyresource exists, use a different cluster for the DRA workflow, or remove the existing installation first. Refer to Cleanup.
Note
OpenShift’s CRI-O container runtime supports the Container Device Interface (CDI) required by the DRA driver without additional configuration.
Installing the GPU Operator#
Follow these steps to install the NVIDIA GPU Operator by using the OpenShift CLI (oc), then create GPUCluster and NVIDIADriver resources instead of ClusterPolicy.
Create a namespace for the NVIDIA GPU Operator and save it in the
nvidia-gpu-operator.yamlfile:apiVersion: v1 kind: Namespace metadata: name: nvidia-gpu-operator
$ oc create -f nvidia-gpu-operator.yamlCreate an
OperatorGroupCR and save it in thenvidia-gpu-operatorgroup.yamlfile:apiVersion: operators.coreos.com/v1 kind: OperatorGroup metadata: name: nvidia-gpu-operator-group namespace: nvidia-gpu-operator spec: targetNamespaces: - nvidia-gpu-operator
$ oc create -f nvidia-gpu-operatorgroup.yamlNote
Refer to Installing the NVIDIA GPU Operator using the CLI for detailed namespace and
OperatorGroupcreation steps if you have not already created them.Select the
v26.7channel explicitly and look up the starting CSV. DRA support requires v26.7 or later, and the certified-operator default channel might still point to an earlier release:$ CHANNEL=v26.7 $ STARTING_CSV=$(oc get packagemanifests/gpu-operator-certified -n openshift-marketplace -ojson | jq -r '.status.channels[] | select(.name == "'$CHANNEL'") | .currentCSV')
Create the
SubscriptionCR:$ cat <<EOF > nvidia-gpu-sub.yaml apiVersion: operators.coreos.com/v1alpha1 kind: Subscription metadata: name: gpu-operator-certified namespace: nvidia-gpu-operator spec: channel: $CHANNEL installPlanApproval: Manual name: gpu-operator-certified source: certified-operators sourceNamespace: openshift-marketplace startingCSV: $STARTING_CSV EOF $ oc create -f nvidia-gpu-sub.yaml
Approve the install plan:
$ INSTALL_PLAN=$(oc get installplan -n nvidia-gpu-operator -oname) $ oc patch $INSTALL_PLAN -n nvidia-gpu-operator --type merge --patch '{"spec":{"approved":true }}'
Create the GPUCluster instance#
The GPUCluster custom resource definition (CRD) and a default example are provided by the GPU Operator CSV, the same way ClusterPolicy is provided.
Extract the default
GPUClusterexample from the CSV:$ oc get csv -n nvidia-gpu-operator $STARTING_CSV -o jsonpath='{.metadata.annotations.alm-examples}' | jq -r 'map(select(.kind == "GPUCluster")) | .[0]' > gpucluster.json
The default example is similar to the following:
{ "apiVersion": "nvidia.com/v1alpha1", "kind": "GPUCluster", "metadata": { "name": "gpu-cluster" }, "spec": { "draDriver": { "repository": "nvcr.io/nvidia", "image": "dra-driver-nvidia-gpu", "version": "v0.5.0", "imagePullPolicy": "IfNotPresent", "computeDomains": { "enabled": true } }, "dcgm": { "enabled": false }, "dcgmExporter": { "enabled": true } } }
Note
GPUClusterdoes not manage the NVIDIA GPU driver. Create anNVIDIADriverresource in the next section to install the driver.Apply the
GPUClusterresource:$ oc apply -f gpucluster.jsongpucluster.nvidia.com/gpu-cluster created
Create the NVIDIADriver instance#
Extract the default
NVIDIADriverexample from the CSV:$ oc get csv -n nvidia-gpu-operator $STARTING_CSV -o jsonpath='{.metadata.annotations.alm-examples}' | jq -r 'map(select(.kind == "NVIDIADriver")) | .[0]' > nvidiadriver.json
The default example is similar to the following (abbreviated):
{ "apiVersion": "nvidia.com/v1alpha1", "kind": "NVIDIADriver", "metadata": { "name": "gpu-driver" }, "spec": { "driverType": "gpu", "repository": "nvcr.io/nvidia", "image": "driver", "version": "sha256:<digest>", "nodeSelector": {} } }
Note
The default
versionfield pins a specific driver build by image digest rather than a release tag such as580.65.06. The digest satisfies the DRA driver’s minimum required GPU driver version of 580 or later. Changerepository,image, andversionto reference a different driver image if required, following the same pattern as theClusterPolicydriverfields described in Create the ClusterPolicy instance.Apply the
NVIDIADriverresource:$ oc apply -f nvidiadriver.jsonnvidiadriver.nvidia.com/gpu-driver createdThe Operator builds the driver container by using the OpenShift Driver Toolkit (DTK), the same mechanism used for the
ClusterPolicydriver daemonset. If the driver pod does not become ready, refer to About the Broken Driver Toolkit for troubleshooting steps that apply to both workflows.
Verify the installation#
Confirm that the
GPUClusterandNVIDIADriverresources are ready:$ oc get gpucluster gpu-cluster $ oc get nvidiadriver -n nvidia-gpu-operator
Example Output
NAME STATUS AGE gpu-cluster ready 6m38s NAME STATUS DEFAULT AGE gpu-driver ready false 2026-09-30T13:07:35Z
Note
A
readystatus confirms that the Operator reconciled the desired state. If your cluster has no GPU nodes labeled yet, the managed pods in the next step do not start until the Node Feature Discovery Operator labels a GPU node.Confirm that the managed pods are running in the GPU Operator namespace:
$ oc get pods -n nvidia-gpu-operatorExample Output
NAME READY STATUS RESTARTS AGE gpu-operator-6b4b79979d-vl9pw 1/1 Running 0 3h21m nvidia-dcgm-exporter-dra-pc85w 1/1 Running 0 8m45s nvidia-dra-driver-controller-6fc968db58-fzfw7 1/1 Running 0 171m nvidia-dra-driver-kubelet-plugin-dhjcd 2/2 Running 0 8m34s nvidia-dra-validator-thqnh 1/1 Running 0 8m45s nvidia-gpu-driver-rhel9-59f4c8db85-pn5c5 2/2 Running 0 8m45s
The driver
DaemonSetis namednvidia-gpu-driver-rhel9-<hash>and its node selector includes the RHCOSOSTREE_VERSIONlabel, similar to thenvidia-driver-daemonset-<RHCOS-version>naming used byClusterPolicy.Confirm that the DRA
DeviceClassobjects are available:$ oc get deviceclassExample Output
NAME AGE compute-domain-daemon.nvidia.com 3h compute-domain-default-channel.nvidia.com 3h gpu.nvidia.com 3h mig.nvidia.com 3h vfio.gpu.nvidia.com 3h
Confirm that a
ResourceSlicewas published for your GPU node:$ oc get resourceslice -o yamlExample Output (abbreviated)
apiVersion: resource.k8s.io/v1 kind: ResourceSlice spec: devices: - attributes: addressingMode: string: HMM architecture: string: Turing brand: string: Nvidia cudaComputeCapability: version: 7.5.0 cudaDriverVersion: version: 13.2.0 driverVersion: version: 595.91.7 productName: string: Tesla T4 uuid: string: GPU-35cc335d-cbdd-999f-6289-38d96976de78 capacity: memory: value: 15Gi name: gpu-0 driver: gpu.nvidia.com
Running a sample DRA workload#
Request a full GPU by creating a ResourceClaimTemplate and a pod that references it.
Create a project for the sample workload:
$ oc new-project gpu-exampleCreate a
ResourceClaimTemplatethat requests one GPU by using the managedgpu.nvidia.comDeviceClass:$ cat <<EOF | oc create -f - apiVersion: resource.k8s.io/v1 kind: ResourceClaimTemplate metadata: namespace: gpu-example name: single-gpu spec: spec: devices: requests: - name: gpu exactly: deviceClassName: gpu.nvidia.com EOF
Create a pod that references the
ResourceClaimTemplate:$ cat <<EOF | oc create -f - apiVersion: v1 kind: Pod metadata: namespace: gpu-example name: gpu-pod spec: containers: - name: workload image: nvcr.io/nvidia/k8s/cuda-sample:vectoradd-cuda12.5.0-ubi8 command: ["bash", "-c"] args: ["nvidia-smi -L; sleep 9999"] resources: claims: - name: gpu resourceClaims: - name: gpu resourceClaimTemplateName: single-gpu EOF
Note
OpenShift can report a
PodSecurityadmission warning similar to the following. The warning does not block pod creation under the default namespace Pod Security level and can be ignored for this example:Warning: would violate PodSecurity "restricted:latest": allowPrivilegeEscalation != false (container "workload" must set securityContext.allowPrivilegeEscalation=false), unrestricted capabilities (container "workload" must set securityContext.capabilities.drop=["ALL"]), runAsNonRoot != true (pod or container "workload" must set securityContext.runAsNonRoot=true), seccompProfile (pod or container "workload" must set securityContext.seccompProfile.type to "RuntimeDefault" or "Localhost")Verify that the pod was allocated a GPU:
$ oc exec -n gpu-example gpu-pod -- nvidia-smi -L
Example Output
GPU 0: Tesla T4 (UUID: GPU-35cc335d-cbdd-999f-6289-38d96976de78)Clean up the sample workload:
$ oc delete project gpu-example
For more allocation patterns, such as selecting a GPU by product name or memory size, sharing a GPU across containers, or requesting multiple GPUs, refer to Request full GPUs in the upstream DRA driver documentation.
Multi-Node NVLink and ComputeDomain#
By default, the GPUCluster resource enables ComputeDomain support (draDriver.computeDomains.enabled: true), which deploys the ComputeDomain controller and kubelet plugin and publishes generic daemon and channel devices under the compute-domain.nvidia.com driver, even on hardware without Multi-Node NVLink (MNNVL). This is expected and does not require any additional configuration on typical GPU nodes.
ComputeDomain support is intended for NVIDIA Grace Blackwell systems with Multi-Node NVLink, such as NVIDIA HGX GB200 NVL72 or NVIDIA HGX GB300 NVL72. Configuring and using ComputeDomains for multi-node workloads is out of scope for this guide. Refer to DRA Driver for NVIDIA GPUs for ComputeDomain prerequisites and configuration, and to GPU Operator with KubeVirt and DRA for KubeVirt and OpenShift Virtualization VFIO passthrough with DRA.
Uninstall#
Delete any user-created
ResourceClaimsand workload pods that reference GPUs, and confirm that the claims are removed.Delete the
NVIDIADriverandGPUClusterresources:$ oc delete nvidiadriver -n nvidia-gpu-operator --all $ oc delete gpucluster gpu-cluster
Uninstall the GPU Operator subscription and CSV. Refer to Cleanup for the complete uninstall procedure.
For finalizer-related recovery procedures, such as recovering from an uninstall that leaves orphaned ResourceClaims, refer to the Troubleshooting section of DRA Driver for NVIDIA GPUs.