Deploy GPU Operator in an IPv6 Network#
Introduction#
You can deploy the NVIDIA GPU Operator on an OpenShift Container Platform cluster that uses IPv6-only networking. The only component that requires IPv6-specific configuration is NVIDIA DCGM.
The default ClusterPolicy custom resource sets dcgm.enabled to true.
With this setting, the Operator deploys DCGM as a separate nvidia-dcgm daemon set.
DCGM Exporter connects to the DCGM host engine, nv-hostengine, through the nvidia-dcgm service on port 5555.
The DCGM container starts the host engine with the -b 0.0.0.0 argument by default.
This argument binds the host engine to IPv4 addresses only.
In an IPv6-only cluster, DCGM Exporter cannot connect to the host engine and the exporter pods go into CrashLoopBackOff.
Configure DCGM for IPv6#
To bind the host engine to IPv6 addresses, set dcgm.args and specify :: as the bind address.
The :: value binds the host engine to all interfaces for IPv4 and IPv6.
The values in dcgm.args replace the default container arguments, so you must specify the complete argument list.
Prerequisites
GPU Operator v26.7.1 or later. Earlier versions include a DCGM Exporter that cannot resolve the
nvidia-dcgmservice name to an IPv6 address.
Procedure
If you create the cluster policy from the command line, add the following
dcgmfield to thespecobject in theclusterpolicy.jsonfile before you apply it:"dcgm": { "enabled": true, "args": ["-n", "-b", "::", "--log-level", "NONE", "-f", "-"] }
If the cluster policy already exists, patch it:
$ oc patch clusterpolicy gpu-cluster-policy --type=merge \ -p '{"spec":{"dcgm":{"args":["-n","-b","::","--log-level","NONE","-f","-"]}}}'