Quickstart#
Set up with a coding agent#
The fastest path: paste this to your agent and it does the rest, including walking you through the choices below. Refer to Skills for how it works.
Set up xr-ai for me: ask me whether to build against the latest release
(newest stable tag; a prerelease only if no stable exists) or main, then
fetch skills/getting-started/SKILL.md at that ref from
https://raw.githubusercontent.com/NVIDIA/xr-ai/. If the file does not exist
at that ref, the release predates agent setup: get my OK to use main for
both the skill and the checkout. Install the skill (or just follow it), and
walk me through the rest of the setup.
The remainder of this quickstart is the manual path.
Every sample follows the same pattern: start the shared model stack, wait for
its launcher to report readiness and return, start the sample from the same
terminal, then connect a client. Once the sample is ready, any supported
client — web browser, Android app, iOS/visionOS app, or AR glasses — can join
the session using the token printed on startup. Each procedure below starts
with a cd from the repository root; keep running that procedure’s commands
from the selected sample directory.
Each interactive procedure includes one concrete spoken request that exercises the sample’s main path. After it succeeds, you can ask “What can you do?” to explore further. In samples with wake-word gating, prefix that spoken question with “Hey Agent”; one follow-up utterance within five seconds can omit the wake phrase. Typed requests bypass the voice gate. Treat the reply as a conversational introduction rather than an exhaustive capability reference.
Simple VLM example (vision Q&A over voice + text)#
End-to-end voice + vision sample. Speak into the mic or type into the data
channel; both routes use the same VLM pipeline against the latest video frame.
Replies arrive as streaming Pocket TTS audio plus a vlm.response text message.
Uses the text-output Reasoner from nvidia/Cosmos3-Nano by default. Refer to
AI services for runtime-selection details.
The sample always reuses model services and never starts or stops them. Its
fixed yaml/models.json expects Parakeet STT on port 8103, Cosmos3-Nano on
port 8100, and Pocket TTS on port 8105. From the sample directory, start the
repository defaults first:
cd agent-samples/simple-vlm-example
uv run --project ../../model-server-samples/model-servers model_servers
The command may download model weights on its first run. Refer to the credentials guide for the required credentials. The model services remain running across sample restarts.
Step 1 — Start the server#
uv sync
uv run simple_vlm_example
Alternatively, run the source file directly after synchronization:
uv run main.py
Only DeviceIOHub and the worker start. For the worker’s startup warmup, refer to Simple VLM example. For readiness and required deployment order, refer to Consumer model readiness. The hub prints:
[hub] LiveKit URL : wss://localhost:8080
[hub] Room : xr-room
[hub] Token : eyJ…
[hub] Web client : https://localhost:8080
This banner appears as soon as the hub itself is ready, while the model
services and worker are still starting. Clients can connect as soon as it
appears, but the agent answers queries only after the launcher prints its
All processes ready banner.
Step 2 — Connect a client#
Open https://localhost:8080 in a browser. The samples ship with HTTPS on by
default (a development root CA and signed server leaf are generated on first
run under ~/.local/share/xr-ai/), so you’ll see a “Your connection is not
private” warning the first time — click Advanced → Proceed (Chrome or Edge) or
Accept the Risk and Continue (Firefox). Refer to the networking guide for
trusting the certificate permanently or running over plain HTTP instead.
Leave Token URL blank — the web client fetches a token from the server automatically. Click Connect, then click Start Camera and allow camera access. For a spoken interaction, click Start Microphone and allow microphone access.
Step 3 — Meet the agent#
Try this first spoken interaction:
You: Hey Agent, what am I looking at?
Agent: [Answers from the current camera frame.]
Type any question → sent verbatim to the VLM.
Speak into your mic → speech is transcribed and sent as a query.
A successful round trip: your query appears in the log, the agent responds after a moment, and you hear the reply through your speakers.
To use compatible services at different locations, edit their endpoints in
yaml/models.json:
{
"endpoint": {"base_url": "https://your-vlm.example.com"}
}
The sample does not offer deployment profiles; the referenced services remain operator-owned regardless of endpoint.
Each sample has its own device_io_hub.yaml controlling the hub; refer to
services/device-io-hub/device_io_hub.yaml for the full option list.
Lab instrument monitoring (marker-associated readings + foreground voice)#
This sample offers one on-demand background visual observation task per
participant while a separate generic tool-calling agent answers voice or typed
queries and controls that task. A separate QR and ArUco instrument monitor tracks readings,
speaks only discovered, changed, or long-missing device updates, and persists
10-second full-state snapshots. The sample writes monitor, instrument,
final pre-gate transcript, foreground-turn, and Relay JSONL files under artifacts/
and intentionally serves no sample-specific monitoring web UI. Its fixed model
configuration uses Cosmos for visual inference.
Start the shared model-servers stack, which includes Pocket TTS:
cd agent-samples/lab-instrument-monitoring
uv run --project ../../model-server-samples/model-servers model_servers
Wait for the model launcher to report readiness and return, then start the sample from the same terminal:
uv sync
uv run lab_instrument_monitoring
Alternatively, run the source file directly after synchronization:
uv run main.py
Connect an existing glasses or platform client using the authenticated URL, room, and token printed by the hub. The token and signaling routes remain available on port 8080 together with the shared connection web client. Refer to the lab instrument architecture guide for reusable agent patterns, marker setup, output contracts, and adaptation recipes.
Try the lab agent#
Point the camera at instruments carrying the configured markers, then try this spoken interaction:
You: Hey Agent, read the instruments.
Agent: [Reports the visible marker-associated readings or their availability.]
Inspect the JSONL artifacts while trying additional interactions.
Tea-making guidance (voice + visual workflow)#
This sample combines an interactive tea guide with optional background change,
transcript, and video observation. Nemotron-3 Nano Omni supplies both language
and visual inference. Records are written as JSON Lines under the sample’s
artifacts/ directory. A separate live event viewer presents selected runtime
events without replacing those durable records.
Start the shared model services, including Pocket TTS:
cd agent-samples/tea-making-sample
uv run --project ../../model-server-samples/model-servers model_servers
Wait for the model launcher to report readiness and return, then launch the sample from the same terminal:
uv sync
uv run tea_making_sample
Alternatively, run the source file directly after synchronization:
uv run main.py
Open the DeviceIOHub connection page at https://localhost:8080, accept the
development certificate on first use, allow camera and microphone access, and
connect. The checked-in voice-gate YAML requires “Agent” or “Hey Agent.” Set
voice_gate_yaml: voice_gate.always-on.yaml in yaml/tea_making_worker.yaml
to dispatch every finalized utterance without a wake phrase.
Open http://127.0.0.1:8092 on the XR-AI host for the live event viewer. To
view it directly from another trusted machine, add --expose-web-events and
use http://<xr-host>:8092; restrict that unauthenticated port to the trusted
development network.
The tea workflow changes steps only after an explicit user command. Visual observations can satisfy the current step’s evidence requirements, but never advance the workflow silently. Refer to the tea-making architecture guide for reusable workflow patterns, background-agent contracts, backend integration, and adaptation recipes.
Try the tea agent#
Try this first spoken interaction:
You: Hey Agent, help me make tea.
Agent: [Starts the tea guide and presents the first step.]
XR render demo (voice-driven sphere in CloudXR)#
Speak to the web client and a sphere in the streamed scene tracks your voice — radius follows loudness, colour and position follow spoken commands (“make it red”, “put it to my left”, “where I’m looking”). Runs against a Quest 3 or Vision Pro on the same LAN, or the IWER emulator built into the web client for desktop dev.
Under the hood, the orchestrator launches the hub, CloudXR runtime, typed capability processes, and the worker alongside the reused model endpoints. The worker calls those processes through Relay-managed native tools. The voice runtime runs quick-acks and a Nemotron-3-Nano-Omni-30B-A3B-Reasoning agentic tool-calling loop over scene, tracking, spatial math, vision, and video-memory tools. Refer to the xr-render-demo reference for the full process map, agentic-loop details, and XR session lifecycle.
Requires model-servers to be running first — the demo does not start its
own model services.
Step 1 — Start model servers (once)#
cd agent-samples/xr-render-demo
uv run --project ../../model-server-samples/model-servers model_servers
This exits immediately once all configured services are ready. Weights stay loaded in the background.
Step 2 — Start the demo#
Refer to the shared Requirements first. This demo has two additional host prerequisites:
Vulkan loader + headers — the CloudXR compositor and LOVR render through Vulkan, so install them before running the demo:
sudo apt install libvulkan-devNode.js 20.19.0+ with npm on PATH for WebRTC profiles — the orchestrator builds the web vendor bundle on first run (skipped for native profiles and on subsequent runs).
Start XR Render:
uv sync
uv run xr_render_demo
Alternatively, run the source file directly after synchronization:
uv run main.py
On first run the orchestrator automatically downloads the pinned LOVR version to
deps/lovr/ inside the repository. For WebRTC profiles it also builds the web
vendor bundle, which requires npm and network access. Existing artifacts are
reused on subsequent runs.
Note
On DGX Spark (aarch64), LOVR does not publish a prebuilt aarch64 Linux
binary, so the auto-download is not available: build LOVR from source and export
LOVR_BIN. Refer to Troubleshooting for build instructions.
To use a custom LOVR build:
export LOVR_BIN=/path/to/your/lovr # or set lovr_bin: in scene/scene_service.yaml
uv run xr_render_demo
GPU pinning for the XR side is controlled by gpu_index in
yaml/cloudxr_runtime.yaml. cloudxr-runtime applies
the pin to its own process and writes the selectors into cloudxr.env;
the scene process and LOVR inherit from that file. Refer to the
xr-render-demo reference for full details.
Step 3 — Connect and meet the agent#
Open the authenticated web-client URL printed by DeviceIOHub. Accept the port-8080 development certificate on first use, leave Token URL blank, and click Connect. Wait for the agent status to report ready, then click Start Microphone. In the XR Stream section, follow the accept the CloudXR cert link to accept the second development certificate, return to the client, and click Launch XR.
With the default always-on setting, ordinary finalized speech is dispatched without a wake phrase; stop commands are intercepted by the voice gate. Try this first spoken interaction:
You: Make the sphere red.
Agent: [Changes the sphere's color or explains why it could not.]
Speak naturally: the agent accepts questions, commands, and follow-up requests. The scene draws two shapes, boxes and spheres; ask for either by name. Refer to Connecting clients for headset and native-client setup.
To stop the model servers when done:
uv run --project ../../model-server-samples/model-servers model_servers --stop
XR Render uses the configured shared endpoints in yaml/models.json; it does
not select or own model deployment profiles.
Hub only (standalone)#
cd services/device-io-hub
uv sync
uv run device_io_hub --config device_io_hub.yaml
Useful for development or when running an agent in a separate terminal. The explicit configuration is the repository reference copy with every field documented beside its value.