Skip to content

AI Skills

The repository includes task-specific SKILL.md playbooks that help compatible AI agents operate l8k consistently. They describe supported commands, flags, safety boundaries, structured output, and troubleshooting workflows.

AI skills are documentation for an agent. They are not Launch Kit runtime plugins and do not change the l8k binary.

Skill Catalog

Skill Agent task
k8s-network-engineer Route broad NVIDIA Kubernetes networking requests to the appropriate workflow.
k8s-launch-kit-shared Apply common install paths, output rules, error handling, and safety guidance.
k8s-launch-kit-config Create, inspect, and edit Launch Kit configuration.
k8s-launch-kit-discover Inventory cluster hardware and produce cluster-config.yaml.
k8s-launch-kit-generate Select a profile and render manifests.
k8s-launch-kit-dryrun Preview generated or server-side deployment changes.
k8s-launch-kit-deploy Apply generated resources in dependency order.
k8s-launch-kit-clean Remove Network Operator custom resources and uninstall or retain its Helm release according to ownership config.
k8s-launch-kit-validate Run deployment acceptance and interpret the validation report.
k8s-launch-kit-pipeline Coordinate discovery, generation, and deployment as an end-to-end flow.
k8s-launch-kit-troubleshoot Diagnose discovery, operator, SR-IOV, RDMA, and IPAM failures.

The source playbooks are under skills/. Each phase skill depends on the shared skill for common behavior.

Make Skills Available

Use the project or workspace skill mechanism provided by the AI agent. Point it at the required skill directories, preserving each directory and its bundled references/ files. Installation and discovery locations are agent-specific.

For a complete operational assistant, expose k8s-network-engineer and all of its required k8s-launch-kit-* skills. For a narrower automation task, expose k8s-launch-kit-shared plus only the phase skills the agent needs.

Clone the repository before creating links:

git clone https://github.com/NVIDIA/k8s-launch-kit.git
cd k8s-launch-kit

Common agent layouts:

mkdir -p ~/.claude/skills
for skill in skills/k8s-launch-kit-* skills/k8s-network-engineer; do
  ln -sfn "$PWD/$skill" "$HOME/.claude/skills/$(basename "$skill")"
done

Use <project>/.claude/skills/ instead for project-scoped skills.

mkdir -p .agents/skills
for skill in skills/k8s-launch-kit-* skills/k8s-network-engineer; do
  ln -sfn "$PWD/$skill" ".agents/skills/$(basename "$skill")"
done

Use ~/.agents/skills/ for user-scoped skills. An AGENTS.md can still carry project-specific policy that applies alongside the skills.

mkdir -p .cursor/rules
for skill in skills/k8s-launch-kit-* skills/k8s-network-engineer; do
  name=$(basename "$skill")
  cp "$skill/SKILL.md" ".cursor/rules/${name}.mdc"
done

Copy or expose any referenced files as project context when the selected rule points to references/.

For other agents, load the relevant SKILL.md files as persistent project context or expose the skills/ tree through the agent's resource mechanism. YAML frontmatter is metadata; agents that do not parse it can still use the Markdown body.

Example request after the skills are available:

Discover this cluster, review the generated profile, render the manifests,
perform a server-side dry run, and stop before deployment.

Machine-Readable CLI Contract

Skills complement the CLI's structured interface. They instruct agents to inspect capabilities rather than infer flags:

l8k schema | jq .

Lifecycle JSON mode sends human-readable logs to stderr. Root, discover, and generate use one result envelope; validation emits a stream. Standalone deploy has no finalized success envelope, and preset/sosreport success remains text. Use the command-specific contract and preserve the process exit status before parsing:

l8k discover \
  --save-cluster-config ./cluster-config.yaml \
  --output json >discover.json 2>discover.log

l8k generate \
  --user-config ./cluster-config.yaml \
  --save-deployment-files ./deployment \
  --output json >generate.json 2>generate.log

Check each command's exit status before parsing its result, and retain stderr for diagnosis.

Do not add --yes to subcommands. --output json is the portable non-interactive path for those commands.

Deployment Boundaries

An AI agent should:

  • Reuse the profile persisted by discovery unless the user requests an override.
  • Generate and inspect manifests before applying them.
  • Use l8k deploy --dry-run for a server-side preview.
  • Require clear authority before changing a live cluster.
  • Treat cleanup as destructive: verify the kubeconfig and resolved operator namespace before running l8k clean.
  • Treat --overwrite-existing as an explicit decision after reviewing Helm drift and every stray deletion, including resources without l8k ownership annotations.
  • Run l8k validate as the normal acceptance stage after deployment.
  • Use l8k sosreport and focused Kubernetes inspection when acceptance fails.

The skill files provide procedure and guardrails; Kubernetes credentials and authorization remain external to the skill.

Example Agent Tasks

Discovery and review:

Discover the cluster, summarize every source group, explain any preset
deviations, and stop before generating manifests.

Render without applying:

Generate an SR-IOV Ethernet deployment for all H200 groups, inspect the
Helm and Kubernetes diffs with a server-side dry run, and do not deploy.

Acceptance and triage:

Run the normal validation workflow. If it does not pass, retain the test
DaemonSet, collect a sosreport, and identify the first failed stage.

Developing Launch Kit

For source changes, start with the repository's AGENTS.md. It covers development workflow, package boundaries, verification and required documentation updates. CLAUDE.md points to that shared guide.

For extension work, consult the contract index. It defines shared integration behavior and scenarios for configuration, discovery, lifecycle, platform, profile, artifacts, resource kinds, ownership, connectivity, and output. Read the applicable baseline exceptions; the specs do not imply that every scenario is already enforced. Update a contract when its shared promise changes, and update the affected user guide and skill alongside implementation. Routine conforming additions do not need a new spec or proposal.

Use operational skills from the checkout being changed when testing CLI workflows. Source development can use make build and ./build/l8k from the repository root; a global installation and cluster access are not prerequisites.

Maintaining Skills

Update relevant documentation sections, the corresponding skill and its bundled references in the same PR whenever a CLI workflow, flag, default, exit code, or safety requirement changes. Keep examples aligned with l8k schema and command help, and keep shared behavior in k8s-launch-kit-shared instead of duplicating it across every phase.