Install GPU Operator in IPv6 Networks#
About IPv6-Only Clusters#
You do not need any IPv6-specific configuration to install the Operator in an IPv6-only cluster.
The only exception is when you deploy DCGM separately from DCGM Exporter.
By default, DCGM Exporter runs an embedded DCGM host engine, nv-hostengine, inside the exporter container.
If you set dcgm.enabled to true, the Operator deploys DCGM as a separate nvidia-dcgm daemon set.
DCGM Exporter then connects to the host engine through the nvidia-dcgm service on port 5555.
In an IPv6-only cluster, you must configure the separate host engine to listen on IPv6 addresses.
Configure Standalone DCGM for IPv6 Clusters#
The DCGM container starts the host engine with the -b 0.0.0.0 argument by default.
This argument binds the host engine to IPv4 addresses only.
In an IPv6-only cluster, DCGM Exporter cannot connect to the host engine and the exporter pods go into CrashLoopBackOff.
To bind the host engine to IPv6 addresses, set dcgm.args and specify :: as the bind address.
The :: value binds the host engine to all interfaces for IPv4 and IPv6.
The values in dcgm.args replace the default container arguments, so you must specify the complete argument list.
Prerequisites
GPU Operator v26.7.1 or later. Earlier versions include a DCGM Exporter that cannot resolve the
nvidia-dcgmservice name to an IPv6 address.
Procedure
Create a file, such as
dcgm-values.yaml, with the following content:dcgm: enabled: true args: - "-n" - "-b" - "::" - "--log-level" - "NONE" - "-f" - "-"
Install the Operator with the values file:
$ helm install --wait --generate-name \ -n gpu-operator --create-namespace \ nvidia/gpu-operator \ --version=v26.7.1 \ -f dcgm-values.yaml
If the Operator is already installed, you can patch the cluster policy instead:
$ kubectl patch clusterpolicies.nvidia.com/cluster-policy --type=merge \ -p '{"spec":{"dcgm":{"enabled":true,"args":["-n","-b","::","--log-level","NONE","-f","-"]}}}'