Skip to content

Set Up NRI Injection

Mokka's NRI plugin is how a workload that was given GPUs the usual way runs nvidia-smi and loads the mock NVML library, with no Mokka-specific pod spec changes. The chart runs it by default. This page covers what the node needs, how to check that the plugin injects, and how to keep workloads out.

What the plugin does

The plugin runs as the nvml-mock-nri container in the node DaemonSet. As each container is created, it adds the mock driver to the ones that should have it:

  • A container holding a GPU allocation from the NVIDIA device plugin or the NVIDIA DRA driver gets the mock driver and keeps exactly its allocated GPUs.
  • A pod annotated nvml-mock.nvidia.com/devices: "true", such as a node-wide monitoring agent, gets every mock GPU on the node without a GPU request.
  • InfiniBand and IMEX channels are selected separately, each by its own annotation.
  • Every other container is left exactly as authored.

Which containers are injected holds the full rules, and Annotations lists every annotation.

Prerequisites

  • containerd 1.7 or later, with NRI enabled on every GPU node. containerd 2.0 and later enable it by default; containerd 1.7 does not. Check a node's effective configuration:

    containerd config dump | grep -A3 'io.containerd.nri.v1.nri'
    #   [plugins.'io.containerd.nri.v1.nri']
    #     disable = false
    #     socket_path = '/var/run/nri/nri.sock'
    #     plugin_path = '/opt/nri/plugins'
    

    That is containerd 2.x. containerd 1.7 prints the same keys with double quotes.

  • For a NVIDIA device plugin, --pass-device-specs=true or the cdi-cri device list strategy, which the GPU Operator uses with CDI enabled. With neither, the allocation leaves nothing in the container the plugin can recognise; see Recognising a GPU allocation.

Current Kind node images ship containerd 2.x and need nothing extra. Older ones run containerd 1.7, for example kindest/node:v1.28.15; create those clusters with NRI enabled in containerd:

kind.yaml
kind: Cluster
apiVersion: kind.x-k8s.io/v1alpha4
containerdConfigPatches:
  - |-
    [plugins."io.containerd.nri.v1.nri"]
      disable = false
      disable_connections = false
      socket_path = "/var/run/nri/nri.sock"
nodes:
  - role: control-plane
  - role: worker
kind create cluster --name mokka-nri --config kind.yaml

Install Mokka

Install Mokka into its own namespace. The plugin never injects its own release namespace or kube-system, so installing into a shared namespace would leave that namespace's workloads without the mock.

helm install nvml-mock oci://ghcr.io/nvidia/k8s-test-infra/chart/nvml-mock \
  --namespace mokka --create-namespace \
  --wait --timeout 180s

An existing release turns NRI on at its next helm upgrade. Upgrade from chart 0.4.0 without --reuse-values, which renders 0.4.0's values, including nri.enabled=false and the 0.4.0 image tag. --reset-then-reuse-values (Helm 3.14+) keeps your own overrides instead.

Each node pod runs the NRI plugin next to the node agent, and is Ready only once the plugin has registered with containerd:

kubectl -n mokka get pods -l app.kubernetes.io/name=nvml-mock \
  -o custom-columns='POD:.metadata.name,NODE:.spec.nodeName,READY:.status.conditions[?(@.type=="Ready")].status,INIT:.spec.initContainers[*].name,CONTAINERS:.spec.containers[*].name'

On Kubernetes 1.29 and later the node agent is listed under INIT, as a restartable init container, and nvml-mock-nri under CONTAINERS; see NRI pod lifecycle. On older clusters, or with nri.nativeSidecar=false, both are under CONTAINERS.

A node pod that stays NotReady, with nvml-mock-nri crash-looping, means containerd on that node has NRI disabled; see Troubleshooting.

Verify injection

Run a pod that opts in with the devices annotation, so the check needs no GPU consumer:

kubectl -n default apply -f - <<'EOF'
apiVersion: v1
kind: Pod
metadata:
  name: nri-check
  annotations:
    nvml-mock.nvidia.com/devices: "true"
spec:
  restartPolicy: Never
  containers:
    - name: app
      image: debian:bookworm-slim
      command: ["sleep", "300"]
EOF

kubectl -n default wait --for=condition=ready pod/nri-check --timeout=120s
kubectl -n default exec nri-check -- nvidia-smi -L

The image is debian because the injected nvidia-smi is a glibc binary: a musl image such as busybox or alpine cannot run it. nvidia-smi -L lists every mock GPU on the node. A pod without the annotation and without a GPU request sees none, as on a real GPU node.

To check the allocation path, schedule a GPU request through the device plugin or a ResourceClaim through the DRA driver on a debian-based image: nvidia-smi -L in that pod lists only the GPUs it was allocated.

kubectl -n default delete pod nri-check

Keep a workload out

To Do
Skip one pod Annotate it nvml-mock.nvidia.com/inject: "false"
Skip a namespace Add it to nri.excludedNamespaces
Turn the plugin off helm upgrade nvml-mock ... -n mokka --reuse-values --set nri.enabled=false

Turning the plugin off affects only containers created afterwards. Running containers keep what they were given until they restart.

When it does not inject

The plugin fails open: if it cannot inject, the container starts without the mock rather than failing. Start with A pod gets no mock GPUs even though NRI is enabled, then NRI plugin failure modes.