Deploy Compute Backend#
Install the unified osmo Helm chart with its compute-only profile to
connect a Kubernetes cluster to an existing OSMO control plane. The release
contains the backend listener and worker but no control-plane services or
databases.
Deployment Architecture#
Prerequisites#
Before continuing:
Deploy an OSMO control plane.
Install
kubectl, Helm,jq, and OpenSSL.Install KAI Scheduler in the compute cluster.
Make the control-plane URL reachable from the compute cluster.
Provide enough CPU, memory, and storage for the compute-plane Pods and the workflows you intend to run. GPU workloads also require GPU nodes and the NVIDIA GPU Operator.
The examples use one context for each cluster and keep workflow Pods in a separate namespace:
$ export CONTROL_CONTEXT="control-context"
$ export CONTROL_NAMESPACE=osmo
$ export COMPUTE_CONTEXT="compute-context"
$ export COMPUTE_NAMESPACE=osmo-compute
$ export WORKLOAD_NAMESPACE=osmo-workflows
This guide uses gb200-01 as the backend name. Use the same backend name and
workload namespace in every step.
Install Cluster Dependencies#
Install KAI Scheduler by following the canonical KAI instructions. Use the common kai-values.yaml behavior file, but do not
apply the converged-cluster kai-selectors.yaml overlay: compute-cluster KAI
components follow the placement policy of that cluster.
Prepare Values and Secrets#
Configure the control plane#
The control plane must define the backend before the compute plane connects.
Its k8s_namespace must exactly match the compute chart’s workload
namespace, at least one pool must reference the backend, and workflow Pods
must be able to reach configuration.service.service_base_url. Merge the
following into the complete values used to manage your control-plane release:
configuration:
service:
service_base_url: https://osmo.example.com
backends:
gb200-01:
k8s_namespace: osmo-workflows
pools:
default:
backend: gb200-01
Provision the backend credential#
Generate a backend credential into a protected temporary file and create its Secret in the control-plane namespace. The commands do not print the token:
$ set -o pipefail
$ TOKEN_FILE=$(mktemp) &&
chmod 600 "$TOKEN_FILE" &&
openssl rand -base64 32 | tr -d '\n=' | tr '/+' '_-' > "$TOKEN_FILE" &&
[ -s "$TOKEN_FILE" ] &&
kubectl --context "$CONTROL_CONTEXT" --namespace "$CONTROL_NAMESPACE" \
create secret generic osmo-gb200-01-backend-token \
--from-file=token="$TOKEN_FILE"
$ BACKEND_TOKEN_STATUS=$?
$ rm -f -- "${TOKEN_FILE:-}" || BACKEND_TOKEN_STATUS=$?
$ test "$BACKEND_TOKEN_STATUS" -eq 0
For production, provision the same Secret through your approved secret manager. As a best practice, use a different token for each backend so that each credential can be rotated or revoked independently. Then add a bootstrap identity that references the user-managed Secret to the control-plane values:
authentication:
bootstrap:
identities:
backend-operator-gb200-01:
enabled: true
username: backend-operator-gb200-01
roles:
- osmo-backend
tokens:
primary:
existingSecret:
name: osmo-gb200-01-backend-token
key: token
Apply the updated control-plane release before installing the compute plane. Keep all existing control-plane values in that Helm upgrade.
Copy the Secret to the compute release namespace without decoding or printing it:
$ kubectl --context "$COMPUTE_CONTEXT" create namespace "$COMPUTE_NAMESPACE" \
--dry-run=client -o yaml | kubectl --context "$COMPUTE_CONTEXT" apply -f -
$ kubectl --context "$CONTROL_CONTEXT" --namespace "$CONTROL_NAMESPACE" \
get secret osmo-gb200-01-backend-token -o json \
| jq --arg namespace "$COMPUTE_NAMESPACE" \
'.metadata = {"name":"osmo-gb200-01-backend-token","namespace":$namespace}' \
| kubectl --context "$COMPUTE_CONTEXT" apply --server-side -f -
Prepare compute-plane values#
Create osmo-compute-values.yaml. Replace osmo.example.com with the
control-plane URL that compute-cluster Pods can reach:
externalUrl: https://osmo.example.com
compute:
backendName: gb200-01
workloadNamespace:
name: osmo-workflows
create: true
authentication:
existingSecret: osmo-gb200-01-backend-token
tokenKey: token
If you manage the workload namespace separately, create it before deployment
and set create: false.
Note
Group templates that create ConfigMaps, custom resources, or other
Kubernetes objects need corresponding permissions under
services.backendWorker.extraRBACRules. See
Required Backend Permissions.
Install OSMO#
Pull the chart so the compute profile always matches the selected chart
version, then install it. The command uses Helm 4’s --wait=legacy strategy.
With Helm 3, replace --wait=legacy with --wait:
$ export OSMO_CHART_VERSION=<chart-version>
$ helm repo add osmo https://helm.ngc.nvidia.com/nvidia/osmo
$ helm repo update osmo
$ helm pull osmo/osmo --version "$OSMO_CHART_VERSION" \
--untar
$ helm --kube-context "$COMPUTE_CONTEXT" upgrade --install osmo-compute \
./osmo \
--namespace "$COMPUTE_NAMESPACE" \
--values osmo/profiles/split-plane-compute.yaml \
--values osmo-compute-values.yaml \
--wait=legacy --timeout 10m
Verify the Deployment#
Confirm that the backend listener and worker Deployments are available:
# Verify the Helm release and compute-plane workloads
$ helm --kube-context "$COMPUTE_CONTEXT" status osmo-compute \
--namespace "$COMPUTE_NAMESPACE"
$ kubectl --context "$COMPUTE_CONTEXT" --namespace "$COMPUTE_NAMESPACE" \
rollout status deployment \
--selector app.kubernetes.io/instance=osmo-compute \
--timeout 10m
$ kubectl --context "$COMPUTE_CONTEXT" --namespace "$COMPUTE_NAMESPACE" \
get deployments,pods \
--selector app.kubernetes.io/instance=osmo-compute
$ kubectl --context "$COMPUTE_CONTEXT" --namespace kai-scheduler wait \
--for=condition=Available deployment --all --timeout=10m
An authenticated OSMO CLI is not required to deploy the backend. Optionally, use it to confirm that the backend and pool are online and submit both CPU verification workflows for end-to-end verification:
# Verify the backend, pools, and resources
$ osmo config show BACKEND gb200-01
$ osmo pool list
$ osmo resource list --pool default
# Verify workflow submission and operation
$ osmo workflow submit deployments/workflows/verify-hello.yaml \
--pool default --format-type json
$ osmo workflow submit deployments/workflows/verify-object-storage.yaml \
--pool default --format-type json
$ OSMO_WORKFLOW_ID=<returned-workflow-id>
$ osmo workflow query "$OSMO_WORKFLOW_ID" --format-type json
For each submission, set OSMO_WORKFLOW_ID to the returned workflow ID and
repeat the query until its status is COMPLETED. A FAILED, CANCELLED,
or timed-out workflow is a validation failure.
Troubleshooting#
Unknown backend#
If the backend listener reports that the backend is not configured, add the
exact value of compute.backendName under configuration.backends in the
control-plane values and apply the control-plane release.
Namespace mismatch#
If registration reports a namespace mismatch, make
configuration.backends.<backend-name>.k8s_namespace identical to
compute.workloadNamespace.name and apply the control-plane release.
Backend authentication error#
Verify that both Secrets contain identical token data without printing the decoded credential:
$ CONTROL_TOKEN=$(kubectl --context "$CONTROL_CONTEXT" \
--namespace "$CONTROL_NAMESPACE" \
get secret osmo-gb200-01-backend-token -o jsonpath='{.data.token}')
$ if [ -n "$CONTROL_TOKEN" ]; then
printf '%s' "$CONTROL_TOKEN" | sha256sum
else
echo "Control-plane Secret has no token data" >&2
fi
$ unset CONTROL_TOKEN
$ COMPUTE_TOKEN=$(kubectl --context "$COMPUTE_CONTEXT" \
--namespace "$COMPUTE_NAMESPACE" \
get secret osmo-gb200-01-backend-token -o jsonpath='{.data.token}')
$ if [ -n "$COMPUTE_TOKEN" ]; then
printf '%s' "$COMPUTE_TOKEN" | sha256sum
else
echo "Compute-plane Secret has no token data" >&2
fi
$ unset COMPUTE_TOKEN
If the hashes differ, repeat the Secret-copy step and restart the backend listener and worker.
Connection errors#
From a compute-cluster Pod, verify DNS, TLS trust, firewall rules, and access
to externalUrl. The listener requires a persistent WebSocket connection to
the OSMO gateway.
Workflow remains pending or validation rejects its resources#
Check osmo resource list --pool <pool> and the workflow events. Account for
the CPU and memory requested by OSMO sidecars as well as the user container.
Confirm that KAI Scheduler is running and that the cluster has a node with
enough available capacity for the complete workflow Pod.
Upgrade and Recovery#
Rotate the backend credential#
Use an overlap window so the control and compute planes can change credentials without losing registration:
Update the control-plane Secret so
tokencontains the new value andprevious-tokencontains the old value.Wait for every API replica to accept both credentials.
Replace
tokenin the compute-plane Secret with the new value.Restart the backend-listener and backend-worker Deployments and verify that they reconnect.
Remove
previous-tokenfrom the control-plane Secret.Verify that the old credential is rejected by every API replica.
Cleanup#
Remove the compute release only after its workflows have finished or been canceled. The workload namespace is retained when the chart created it:
$ helm --kube-context "$COMPUTE_CONTEXT" uninstall osmo-compute \
--namespace "$COMPUTE_NAMESPACE" --wait
See also
See /api/configs/backend for backend configuration options and Practical Guide for additional pools and platforms.