Cosmos3-Edge#

This guide covers policy generation and multimodal reasoning with nvidia/Cosmos3-Edge.

Complete Installation first. Policy generation also requires -DBUILD_EXPERIMENTAL_MODELS=ON.

export POLICY_CHECKPOINT=nvidia/Cosmos3-Edge-Policy-DROID
export REASONING_CHECKPOINT=nvidia/Cosmos3-Edge
export ONNX_DIR=$HOME/tensorrt-edgellm-workspace/Cosmos3-Edge/onnx
export ENGINE_DIR=$HOME/tensorrt-edgellm-workspace/Cosmos3-Edge/engines

Policy Generation#

1. Export the Policy on CPU#

tensorrt-edgellm-export \
  "$POLICY_CHECKPOINT" \
  "$ONNX_DIR" \
  --task policy

2. Build All Policy Engines#

./build/experimental_models/cosmos3/examples/cosmos3_policy_build \
  --onnxDir "$ONNX_DIR" \
  --engineDir "$ENGINE_DIR"

Use --maxBatchSize N when the runtime must accept more than one prompt.

3. Run the Policy#

For one observation image:

./build/experimental_models/cosmos3/examples/cosmos3_policy_inference \
  --engineDir "$ENGINE_DIR" \
  --image observation.png \
  --prompt "Pick up the banana and place it in the bowl." \
  --output action.json

For an ordered frame list, replace --image with:

--video frame_00.png,frame_01.png,frame_02.png

The current policy conditions on the most recent frame. The JSON output reports the action tensor, shape [batch, chunk, action_dimension], policy domain, denoise-step count, and whether all values are finite. --steps selects the denoise-step count and --seed controls deterministic initial noise.

Multimodal Reasoning#

1. Export the Reasoner on CPU#

tensorrt-edgellm-export \
  "$REASONING_CHECKPOINT" \
  "$ONNX_DIR/reasoning" \
  --task reasoning

2. Build the LLM and Visual Engines#

./build/examples/llm/llm_build \
  --onnxDir "$ONNX_DIR/reasoning/llm" \
  --engineDir "$ENGINE_DIR/reasoning" \
  --maxInputLen 2048 \
  --maxKVCacheCapacity 4096

./build/examples/multimodal/visual_build \
  --onnxDir "$ONNX_DIR/reasoning/visual" \
  --engineDir "$ENGINE_DIR/reasoning"

3. Run Reasoning#

For image reasoning, create an input file using the standard image message format, then run:

./build/examples/llm/llm_inference \
  --engineDir "$ENGINE_DIR/reasoning" \
  --multimodalEngineDir "$ENGINE_DIR/reasoning" \
  --inputFile input.json \
  --outputFile output.json

For native video-file input, launch the OpenAI-compatible server:

tensorrt-edgellm-serve "$REASONING_CHECKPOINT" \
  --cache-dir "$HOME/tensorrt-edgellm-workspace/cache" \
  --max-image-tokens 4096 \
  --max-image-tokens-per-image 4096 \
  --allowed-local-media-path /data/media

Then send the video itself, rather than converting it to image content blocks:

curl -s http://localhost:8000/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "nvidia/Cosmos3-Edge",
    "messages": [{
      "role": "user",
      "content": [
        {
          "type": "video_url",
          "video_url": {"url": "file:///data/media/example.mp4"},
          "fps": 2.0
        },
        {"type": "text", "text": "Describe the actions in this video."}
      ]
    }],
    "max_tokens": 256
  }'

The server decodes and samples the clip to the visual engine profile and keeps it as one video input through the native runtime.

See the Cosmos3 model guide for the component contracts and implementation details.