Customizing model servers#
Use this workflow to change which shared models run, where they run, or how a sample connects to them. Model-server customization has two separate parts:
A
model-serversdeployment profile declares which shared services the operator starts and owns.Each sample’s active models JSON declares compatible client adapters and endpoints, with deployment ownership set to
reused.
Samples do not start or stop model services. Keep that ownership boundary when adapting a sample: customize the shared stack first, start it, and then point the sample at the resulting endpoints.
Configuration layers#
Layer |
Location |
Responsibility |
|---|---|---|
Deployment profile |
|
Logical roles, client adapters, endpoints, credentials, and shared-service ownership |
Hardware profile |
|
Image or checkpoint, ports, GPU placement, cache paths, and runtime memory settings |
Sample models JSON |
|
The roles that worker consumes and the endpoints it reuses |
Sample worker YAML |
|
Application behavior such as prompts, voice gating, timeouts, and capability-service endpoints |
Do not use --gpu-profile to select models. It selects the reviewed hardware
layout (dual_48G_ada, spark, or 96G_blackwell). The --models option
selects the deployment profile independently.
Start from a shipped deployment profile#
The shipped profiles are:
default: local Parakeet STT, Pocket TTS, Nemotron Omni, Cosmos3-Nano Reasoner, and Nemotron embedding services.vlm_llm_nim: local STT, Pocket TTS, and embedding plus self-hosted Nemotron-3 Nano Omni and Cosmos3-Nano Reasoner NIM containers.
The shared profiles use Pocket TTS. Magpie remains available as a standalone service, but is not integrated into the persistent model stack.
Copy the closest profile under a new name:
cp agent-samples/model-servers/yaml/models.default.json \
agent-samples/model-servers/yaml/models.my-stack.json
Run it by filename stem:
uv run --project agent-samples/model-servers \
model_servers --models my-stack
You can also pass an absolute JSON path or a path relative to the current
working directory to --models.
The profile must retain the wrapped JSON shape:
{
"models": {
"vlm": {
"adapter": {"preset": "cosmos3_nano_reasoner"},
"endpoint": {
"base_url": "http://localhost:8100",
"readiness": "health"
},
"deployment": {
"ownership": "managed",
"service": "vlm"
}
}
}
}
Within the shared profile, managed means model-servers owns that service.
Roles may share a service; for example, llm and agent_llm can both name
omni. A service name must have a corresponding row in _MODEL_SERVICES in
agent-samples/model-servers/main.py.
Customize a hardware-specific server#
Each managed service resolves its YAML from the detected GPU-profile directory. Change the relevant file there when customizing an existing service. Common fields include:
model: nvidia/Cosmos3-Nano
port: 8100
model_cache: ../../../../models
cuda_visible_devices: "0"
gpu_memory_utilization: 0.55
NIM services use image, http_port, nim_cache, and optional NIM_*
environment settings instead:
image: nvcr.io/nim/nvidia/cosmos3-reasoner:1.7.0
http_port: 8100
nim_cache: ../../../../models/nim
cuda_visible_devices: "0"
env:
NIM_MODEL_SIZE: "nano"
NIM_MAX_MODEL_LEN: "16384"
The Cosmos 3 Reasoner container serves Nano as
nvidia/cosmos3-nano-reasoner. Keep the model ID in the deployment profile’s
adapter synchronized with the NIM model size.
When one deployment profile needs a different launch configuration, add a variant beside the base YAML:
nim_vlm_server.yaml
nim_vlm_server_my-stack.yaml
The suffix is the deployment profile filename without models. or .json.
Only add a variant when the base configuration is not valid for that profile.
Review GPU placement and aggregate memory before running several services
together; the checked-in non-Ada NIM settings are estimates where their YAML
comments say so.
Use an endpoint at another address#
If an operator already runs a compatible service elsewhere, no model-server
change is required. Update the sample entry’s endpoint.base_url and keep
ownership reused when the endpoint belongs to an operator-managed XR AI
stack. Use external for a hosted API or another endpoint with no corresponding
launcher process:
{
"endpoint": {
"base_url": "https://integrate.api.nvidia.com",
"api_key_env": "NGC_API_KEY",
"readiness": "none"
},
"deployment": {"ownership": "external"}
}
Do not put the credential value in JSON. Export it or use the credential store. Refer to Credentials for credential options.
Riva speech boundary#
The generic model-server service table and NIM YAMLs still support custom Riva
STT and TTS profiles through the stt-nim and tts-nim service names. No
shipped model-server profile selects them, and the samples intentionally do not
install the optional Riva client. Adding Riva to a sample therefore requires an
explicit worker dependency and models configuration change; it is not an
endpoint-only customization.
Validate and switch profiles#
Before committing a custom profile:
jq empty agent-samples/model-servers/yaml/models.my-stack.json
uv run --project tests pytest -q \
tests/test_model_servers.py \
tests/test_launcher_config.py \
tests/test_nim_docker.py
Model servers persist after the model_servers command reports readiness.
Starting another profile stops persisted services outside the new selection
before launching it. After changing an image, checkpoint, or launch setting,
stop the old stack once so it cannot continue serving stale configuration:
uv run --project agent-samples/model-servers model_servers --stop
uv run --project agent-samples/model-servers \
model_servers --models my-stack
For adapter fields and model capabilities, refer to xr-ai-models. For server runtime behavior, persistence, and NIM credentials, refer to AI inference servers.