Guides¶
Runnable walkthroughs, each one a specific thing you can do with Mokka.
Every guide needs Helm 3.8 or newer and kubectl. Most use Docker and Kind for a local cluster; the Amazon EKS guide can use an existing compatible cluster or create one with Terraform. Each guide states its own target and prerequisites before it installs anything.
Scenarios¶
Standing Mokka up alongside a real consumer. Roughly in order of how much they ask of you — start at the top if you are new.
| Guide | What it shows | Time |
|---|---|---|
| NVIDIA Device Plugin | Mock GPUs advertised as nvidia.com/gpu, and a workload scheduled against them |
~5 min |
| Amazon EKS | The device-plugin path on private, CPU-only managed workers, including the required runtime bootstrap | 20–30 min |
| NVIDIA GPU Operator | The real operator stack — device plugin, GFD, DCGM and the validator — against mock GPUs | ~15 min |
| NVIDIA Dynamo | An OpenAI-compatible endpoint served by Dynamo, its worker scheduled onto a mock GPU through the GPU Operator | ~15 min |
| Slinky (Slurm) | A Slurm cluster whose slurmd discovers mock GPUs as GRES, and srun --gres=gpu:N jobs scheduled against them |
~20 min |
| NVIDIA DRA Driver | Mock GPUs published as ResourceSlices, and a pod scheduled through a ResourceClaim | ~10 min |
| Run:ai fake-gpu-operator | Two node pools — Mokka serving one with a real NVML shim, FGO serving the other | ~10 min |
| ComputeDomain | NVLink fabric identity, with a real nvidia-imex forming a live domain over mock GPUs |
10–20 min |
| NVSentinel | The full health loop: detect a thermal-margin crossing, cordon and drain, then auto-recover on cooldown | ~30 min |
The last two are the most involved: ComputeDomain needs a four-worker cluster with containerd NRI enabled, and NVSentinel pulls the GPU Operator, cert-manager and NVSentinel before it can start.
Tasks¶
Things you do with Mokka, whichever consumer you are running.
| Guide | What it covers |
|---|---|
| Set up NRI injection | Check that pods holding a GPU allocation get the mock driver, and keep workloads out |
| MIG partitioning | Carve a board into MIG slices, advertise each one, and schedule a pod onto a single slice |
| Failure injection | Present a broken GPU — uncorrectable ECC, lost, fallen off the bus — and watch consumers react |
| Use in CI/CD | Run GPU-dependent tests on CPU runners |
| Runtime control | Change temperature, power, utilisation or health on a running node, with no redeploy |
Observability (Prometheus + Grafana)¶
Not a standalone guide. It composes with the GPU Operator rather than replacing
it, so it lives in the Tilt environment instead of shipping its own cluster and
run.sh.
Prometheus scrapes the real, unmodified NVIDIA dcgm-exporter while it reads the
mock libnvidia-ml.so, and Grafana renders the result — on a cluster with no
GPUs. Two manual triggers then inject a temperature or Xid fault and fail if it
never reaches Prometheus, so the scrape path is asserted rather than eyeballed.
See Local Development for the Tilt environment and every flag it takes.