Skip to content

Guides

Runnable walkthroughs, each one a specific thing you can do with Mokka.

Every guide needs Helm 3.8 or newer and kubectl. Most use Docker and Kind for a local cluster; the Amazon EKS guide can use an existing compatible cluster or create one with Terraform. Each guide states its own target and prerequisites before it installs anything.

Scenarios

Standing Mokka up alongside a real consumer. Roughly in order of how much they ask of you — start at the top if you are new.

Guide What it shows Time
NVIDIA Device Plugin Mock GPUs advertised as nvidia.com/gpu, and a workload scheduled against them ~5 min
Amazon EKS The device-plugin path on private, CPU-only managed workers, including the required runtime bootstrap 20–30 min
NVIDIA GPU Operator The real operator stack — device plugin, GFD, DCGM and the validator — against mock GPUs ~15 min
NVIDIA Dynamo An OpenAI-compatible endpoint served by Dynamo, its worker scheduled onto a mock GPU through the GPU Operator ~15 min
Slinky (Slurm) A Slurm cluster whose slurmd discovers mock GPUs as GRES, and srun --gres=gpu:N jobs scheduled against them ~20 min
NVIDIA DRA Driver Mock GPUs published as ResourceSlices, and a pod scheduled through a ResourceClaim ~10 min
Run:ai fake-gpu-operator Two node pools — Mokka serving one with a real NVML shim, FGO serving the other ~10 min
ComputeDomain NVLink fabric identity, with a real nvidia-imex forming a live domain over mock GPUs 10–20 min
NVSentinel The full health loop: detect a thermal-margin crossing, cordon and drain, then auto-recover on cooldown ~30 min

The last two are the most involved: ComputeDomain needs a four-worker cluster with containerd NRI enabled, and NVSentinel pulls the GPU Operator, cert-manager and NVSentinel before it can start.

Tasks

Things you do with Mokka, whichever consumer you are running.

Guide What it covers
Set up NRI injection Check that pods holding a GPU allocation get the mock driver, and keep workloads out
MIG partitioning Carve a board into MIG slices, advertise each one, and schedule a pod onto a single slice
Failure injection Present a broken GPU — uncorrectable ECC, lost, fallen off the bus — and watch consumers react
Use in CI/CD Run GPU-dependent tests on CPU runners
Runtime control Change temperature, power, utilisation or health on a running node, with no redeploy

Observability (Prometheus + Grafana)

Not a standalone guide. It composes with the GPU Operator rather than replacing it, so it lives in the Tilt environment instead of shipping its own cluster and run.sh.

Prometheus scrapes the real, unmodified NVIDIA dcgm-exporter while it reads the mock libnvidia-ml.so, and Grafana renders the result — on a cluster with no GPUs. Two manual triggers then inject a temperature or Xid fault and fail if it never reaches Prometheus, so the scrape path is asserted rather than eyeballed.

make cluster-create
tilt up -- --observability

See Local Development for the Tilt environment and every flag it takes.