Mokka control plane¶
The Stage 1 control plane turns SGPURackProfile and SGPUInventory resources
into control-plane-owned SGPURack resources, assigns eligible Kubernetes Nodes
to logical rack Nodes, and projects the assignment onto those Nodes. It also
decides which SGPURuntimePolicy overrides apply to
each simulated GPU. It models capacity, topology, and runtime state; it does not
provide GPUs to workloads by itself.
Install¶
Install the CRDs, then enable the control plane in the existing chart:
helm upgrade --install mokka-crds deployments/mokka-crds/helm/mokka-crds
helm upgrade --install nvml-mock deployments/nvml-mock/helm/nvml-mock \
--namespace mokka --create-namespace \
--set controlPlane.enabled=true \
--set controlPlane.image.tag=latest
By default the control plane runs ghcr.io/nvidia/mokka-control-plane tagged
with the chart appVersion, which in a released chart is the image published
for the same release. On main the appVersion is the next -dev version,
but main builds of the control-plane image are published only as latest,
so the command above sets controlPlane.image.tag=latest; drop it when
installing a released chart. The controller is cluster-privileged, so pin it
by digest in production: set controlPlane.image.digest to the sha256:...
digest published for the control-plane image. A digest takes precedence over
controlPlane.image.tag.
The control-plane pods are scheduled and labelled independently of the node
DaemonSet, through controlPlane.nodeSelector, controlPlane.tolerations,
controlPlane.affinity, controlPlane.topologySpreadConstraints,
controlPlane.priorityClassName, controlPlane.podLabels and
controlPlane.podAnnotations. They take the same forms as their
node DaemonSet counterparts and the Kubernetes pod
spec. With controlPlane.replicas above 1, pod anti-affinity on
kubernetes.io/hostname keeps the standby replica off the leader's node, and a
spread constraint on topology.kubernetes.io/zone keeps it in another zone:
controlPlane:
replicas: 2
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
topologyKey: kubernetes.io/hostname
labelSelector:
matchLabels:
app.kubernetes.io/name: nvml-mock-control-plane
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: ScheduleAnyway
labelSelector:
matchLabels:
app.kubernetes.io/name: nvml-mock-control-plane
controlPlane.pdb.enabled adds a
PodDisruptionBudget
(PDB) that bounds how many replicas a node drain or cluster upgrade may evict
at once. The PDB is rendered only when controlPlane.replicas is above 1: over
a single replica it would either block every drain or protect nothing. Set
minAvailable or maxUnavailable as a count or a percentage;
maxUnavailable wins if both are set, and with neither minAvailable is 1. A
minAvailable equal to controlPlane.replicas blocks every drain of a node
that runs a replica.
Uninstalling the mokka-crds release retains the CRDs and existing Mokka
resources. Removing the Mokka API and its resources requires deleting the CRDs
explicitly.
Only one Helm release may enable the cluster-wide control plane in a cluster.
Its fixed control-plane.mokka.nvidia.com ClusterRoleBinding is owned by
that release and acts as the singleton guard: another release must wait until
the owner is uninstalled. Upgrades and rollbacks of the owning release keep
using the same guard.
Disable or uninstall safely¶
Do not disable controlPlane.enabled, uninstall its nvml-mock release, or
delete the Mokka CRDs while an SGPUInventory exists. Disabling or uninstalling
removes the control-plane Deployment and RBAC; deleting the CRDs removes its API
objects. Either order prevents the control plane from releasing the
mokka.nvidia.com/inventory-cleanup and mokka.nvidia.com/rack-cleanup
finalizers or removing its Node projection metadata. The CRD chart retains
every CRD on uninstall, so removing either Helm release is not a cleanup
mechanism.
Drain the control plane before removing it. All Mokka custom resources and Nodes
are cluster-scoped; the commands intentionally omit --namespace for them and
discover the namespace of the one owning Helm release from its singleton
ClusterRoleBinding.
MOKKA_RELEASE="$(kubectl get clusterrolebinding \
control-plane.mokka.nvidia.com \
-o jsonpath='{.metadata.annotations.meta\.helm\.sh/release-name}')"
MOKKA_NAMESPACE="$(kubectl get clusterrolebinding \
control-plane.mokka.nvidia.com \
-o jsonpath='{.metadata.annotations.meta\.helm\.sh/release-namespace}')"
test -n "$MOKKA_RELEASE"
test -n "$MOKKA_NAMESPACE"
kubectl --namespace "$MOKKA_NAMESPACE" rollout status deployment \
--selector="app.kubernetes.io/component=control-plane,app.kubernetes.io/instance=$MOKKA_RELEASE" \
--timeout=2m
kubectl delete sgpuinventories.mokka.nvidia.com \
--all --wait=true --timeout=15m
The delete must complete successfully. It waits for the inventory finalizers; those finalizers wait for every control-plane-owned rack and exact Node projection to drain. If it times out, keep the control plane running with its RBAC permissions, inspect the remaining finalizers and control-plane logs, and resolve the reported cleanup conflict. Do not force-remove finalizers.
Verify the drain before changing the Helm release. The first command must
produce no objects. The second must produce no Node names: it checks the two
projection identity markers and any fields still owned by the
mokka-controller server-side apply manager. A clique-only label owned by
another manager is deliberately not treated as Mokka projection state.
kubectl get \
sgpuinventories.mokka.nvidia.com,sgpuracks.mokka.nvidia.com \
--no-headers \
-o 'custom-columns=KIND:.kind,NAME:.metadata.name,FINALIZERS:.metadata.finalizers'
kubectl get nodes --show-managed-fields=true \
-o go-template='{{range .items}}{{$node := .metadata.name}}{{if or (index .metadata.labels "mokka.nvidia.com/sgpu-assigned") (index .metadata.annotations "mokka.nvidia.com/sgpu-assignment")}}{{printf "%s\n" $node}}{{else}}{{range .metadata.managedFields}}{{if eq .manager "mokka-controller"}}{{printf "%s\n" $node}}{{end}}{{end}}{{end}}{{end}}'
Only after both checks are empty may the control plane be disabled or its owning release uninstalled:
# Disable only the control plane; run from the repository root.
helm upgrade "$MOKKA_RELEASE" deployments/nvml-mock/helm/nvml-mock \
--namespace "$MOKKA_NAMESPACE" --reuse-values \
--set controlPlane.enabled=false
# Or uninstall the complete nvml-mock release instead.
helm uninstall "$MOKKA_RELEASE" --namespace "$MOKKA_NAMESPACE"
Profiles, runtime policies, and the retained CRDs may remain for a later
control-plane installation. To remove the API completely, uninstall the
mokka-crds Helm release and then explicitly delete all four retained CRDs,
but only after the drain above:
# Use the namespace where this release was installed.
MOKKA_CRD_NAMESPACE=default
helm uninstall mokka-crds --namespace "$MOKKA_CRD_NAMESPACE"
kubectl delete customresourcedefinitions.apiextensions.k8s.io \
sgpuinventories.mokka.nvidia.com \
sgpuracks.mokka.nvidia.com \
sgpurackprofiles.mokka.nvidia.com \
sgpuruntimepolicies.mokka.nvidia.com
Build the image locally with:
For local Tilt development, pass --control-plane; the control plane and its
CRDs are otherwise disabled and existing Tilt defaults are unchanged.
Declare capacity¶
Apply the example profile and inventory, then make a Node eligible and match the example group selector:
kubectl apply -f examples/controlplane-crds/sgpu-rack-profile.yaml
kubectl apply -f examples/controlplane-crds/sgpu-inventory.yaml
kubectl label node NODE \
mokka.nvidia.com/sgpu-node=true \
mokka.nvidia.com/pool=example
Only Nodes with mokka.nvidia.com/sgpu-node=true enter the control plane's cache.
An empty group selector matches every eligible Node; otherwise both eligibility
and the selector must match. Placement selectors cannot reference
mokka.nvidia.com/sgpu-assigned or nvidia.com/gpu.clique because those labels
are derived from placement. Such an inventory is rejected without changing its
last materialized racks or bindings.
The supported topology envelope is 100,000 eligible Nodes, 100,000 generated
racks, and 64 declared rack groups across all SGPUInventory resources. A rack
group is a homogeneous declaration, not an individual rack: use its count to
expand one group into many racks. Inventories are admitted whole in creation
order while the aggregate limits fit. An inventory outside the envelope reports
CapacityExceeded; deleting or shrinking an older inventory allows the next
declaration to be admitted.
The durable assignment is SGPURack.spec.nodes[].nodeRef. Existing valid
bindings do not move when Nodes, racks, or profiles are added or edited. New
Nodes are ordered by creation time, name, then UID and fill logical Node
coordinates in rack/index order. GPU, serial, fabric, and rack identities derive from the
inventory UID and coordinate, so retries, restarts, and leader changes do not
change an unchanged coordinate.
For each successfully projected binding the control plane owns only:
mokka.nvidia.com/sgpu-assigned=true;nvidia.com/gpu.clique=<fabric UUID>.<clique ID>;mokka.nvidia.com/sgpu-assignment, compact JSON containing exact inventory, rack, profile revision, coordinate, and Node UID data.
Override runtime state¶
An SGPURuntimePolicy overrides the simulated runtime state of some GPUs in one
inventory: their health, device modes, and telemetry. A profile's
defaults.runtime sets the starting state of every GPU the profile renders, and
a policy changes only the fields it sets. An omitted field keeps the value it
would otherwise have, and an explicit zero is a value: drawMilliWatts: 0 sets
the power draw to zero. Stage 1 evaluates policies and reports whether each is
accepted, but it does not apply them to simulated GPUs yet.
The example policy marks the GPU on logical Node 1 of the example rack as failed and drops its power draw to zero:
spec.targetRef names the inventory and can narrow the selection with
rackGroups, rackIndexes, nodeIndexes, and gpuIndexes; an omitted list
selects every index on that axis. The deepest list sets the policy's scope:
Deepest list in targetRef |
Scope |
|---|---|
| none | Inventory |
rackGroups |
Rack group |
rackIndexes |
Rack |
nodeIndexes |
Node |
gpuIndexes |
GPU |
A GPU's effective state starts from its profile defaults and applies the accepted policies that select it from the broadest scope to the narrowest, so a narrower scope overrides a broader one field by field.
Every listed rack group must be declared by the inventory. When a policy selects rack groups of different shapes, each listed index must exist in at least one of them, and each rack group applies only the indexes it has. Node and GPU indexes come from a rack group's profile, so a rack group whose profile is missing has none.
Two policies conflict when they have the same scope, select at least one common
GPU, and set a common field. The policy with the older creationTimestamp stays
in force, with the UID breaking ties, and the other is rejected as a whole.
Deleting the winner accepts the oldest remaining policy.
The Accepted condition reports the outcome:
| Reason | Status | Meaning |
|---|---|---|
Accepted |
True |
The policy applies to its target. |
TargetNotFound |
False |
The target inventory does not exist. |
InvalidTarget |
False |
A listed rack group is not declared, or a listed index selects no GPU. |
Conflicted |
False |
An older policy of the same scope sets one of its fields for some of the same GPUs. |
kubectl get sgpuruntimepolicies shows each policy's target and acceptance, and
the condition message explains a rejection:
kubectl get sgpuruntimepolicies
kubectl get sgpuruntimepolicy example-node-1-failed \
-o jsonpath='{.status.conditions[?(@.type=="Accepted")].message}'
Observe and troubleshoot¶
kubectl get sgpuinventories,sgpuracks,sgpuruntimepolicies
kubectl get sgpuinventory example -o yaml
kubectl get node NODE -o jsonpath='{.metadata.annotations.mokka\.nvidia\.com/sgpu-assignment}'
kubectl -n mokka logs -l app.kubernetes.io/component=control-plane
kubectl -n mokka get lease control-plane.mokka.nvidia.com -o yaml
Inventory conditions distinguish invalid input or profile references,
materialization failures, pending or conflicting placements, and projection
failures. A missing or invalid profile preserves the last materialized racks
but blocks new allocations for that group. Selector overlaps leave an
unassigned Node in conflict. Foreign rack ownership and incompatible Node
metadata are reported and never overwritten. A runtime policy's Accepted
condition names the reason it does not apply.
Shrinking capacity, deleting an inventory or rack, losing eligibility, and
replacing a Node UID remove the exact old projection before clearing a live
binding. Cleanup removes only control-plane-owned keys whose assignment annotation
still names that binding; incompatible values are preserved and retried. If the exact
Node UID no longer exists, cleanup may proceed without touching a same-name
replacement. The singleton ClusterRoleBinding ensures only the owning release
has cluster permissions. Within that release, its namespace-local Lease
prevents replicas from reconciling concurrently; cluster-wide Lease permissions
are unnecessary under the single-installation invariant. Every replica keeps a
synchronized informer view of profiles, inventories, racks, runtime policies,
eligible Nodes, and durable rack-slot assignments. Only the elected leader attaches reconciliation
handlers and runs workers. A standby reports ready after its caches synchronize
and it observes the elected leader; the leader additionally waits for its
handlers to replay the current caches and its workers to start. Upgrades use a
RollingUpdate Deployment strategy with maxUnavailable: 0 and maxSurge: 1
to preserve serving overlap while leader election maintains single-writer mutation.
Stage 1 exclusions¶
Stage 1 has no agent or driver, workload GPU injection, REST API, Redis, heartbeat-based reclamation, Node lease, or cross-rack switch graph. The only Lease is Kubernetes leader election.