pi0.5#

pi0.5 is Physical Intelligence’s flow-matching Vision-Language-Action model. Edge-LLM runs it as four TensorRT components driven by an experimental C++ runtime, which returns an action chunk for a set of camera views and an instruction. See pi0.5 Design for the component contracts and the reasoning behind the split.

This workflow supports three openpi configurations, in FP16, and follows openpi’s policy contract for each. The Hugging Face mirror is a weight distribution channel – its own LeRobot policy configuration is a different contract and is not used. Results are a numerical comparison against openpi’s action chunk, and all three configurations meet the comparator’s gate on their own embodiment’s observations. Task success is not claimed. The exporter consumes an already-converted PyTorch checkpoint (model.safetensors plus config.json); the LeRobot releases are in that form.

--pi05-policy-config

weights

norm stats

state / action

horizon

cameras

pi05_libero

lerobot/pi05_libero_base

physical-intelligence/libero

8 / 7

10

two, both required

pi05_droid

lerobot/pi05_droid

droid

8 / 8

15

exterior and left wrist, both required

pi05_aloha

lerobot/pi05_base

trossen

14 / 14

50

cam_high required, both wrists optional

openpi publishes no pi05_aloha checkpoint: its configuration is an inference contract over the generalist pi05_base weights, which is what the row above names.

Complete Installation first; pi0.5 also requires -DBUILD_EXPERIMENTAL_MODELS=ON. Its action engine carries a trt_edgellm::AttentionPlugin node, so EDGELLM_PLUGIN_PATH must name the built plugin library or that engine cannot be built or deserialized.

export CHECKPOINT=lerobot/pi05_libero_base
export POLICY_CONFIG=pi05_libero
export ONNX_DIR=$HOME/pi05_onnx
export ENGINE_DIR=$HOME/pi05_engines
export BUILD_DIR=/path/to/tensorrt-edge-llm/build
export EDGELLM_PLUGIN_PATH="$BUILD_DIR/libNvInfer_edgellm_plugin.so"

1. Export on CPU#

export PYTHONNOUSERSITE=1
tensorrt-edgellm-export "$CHECKPOINT" "$ONNX_DIR" --dtype float16 \
    --pi05-policy-config "$POLICY_CONFIG"

--pi05-policy-config fixes the action horizon, the camera slots and the prompt form. A converted checkpoint names no configuration, so the flag is required for every one but pi05_libero, whose feature contract names itself. The checkpoint’s own n_action_steps is cross-checked against the configuration, and a pairing whose horizons disagree is refused before anything is built.

This writes one subdirectory per component – visual/, prefix/, action/ and cond/ – plus text_tokenizer/, the token-embedding table, and policy.json, the observation and action contract: shapes from the checkpoint, policy semantics from the openpi configuration. Both the builder and the runtime refuse a bundle without policy.json. The tokenizer is fetched from gs://big_vision/paligemma_tokenizer.model, which openpi loads the same way; the Hugging Face mirror is gated and cannot be fetched unattended.

cond/ precomputes the AdaRMS modulation schedule instead of rebuilding it inside the per-step action graph. The runtime evaluates it only when the denoise-step count or the batch changes, so a run at fixed settings pays it once, not once per request. It is optional: --no-pi05-hoist-adarms-cond exports the other three alone, for debugging and A/B runs.

The LeRobot checkpoints are PyTorch conversions of openpi’s. They carry the processor schemas but not their normalization state – those statistics are policy input/output metadata, not TensorRT calibration data, and they stay in the original openpi checkpoint assets. The runtime refuses to load without them, so fetch the file for the configuration being exported, from the table above:

mkdir -p "$ONNX_DIR/assets"
curl --fail --location --output "$ONNX_DIR/assets/norm_stats.json" \
  https://storage.googleapis.com/openpi-assets/checkpoints/pi05_libero/assets/physical-intelligence/libero/norm_stats.json

# pi05_droid:
#   https://storage.googleapis.com/openpi-assets/checkpoints/pi05_droid/assets/droid/norm_stats.json
# pi05_aloha:
#   https://storage.googleapis.com/openpi-assets/checkpoints/pi05_base/assets/trossen/norm_stats.json

Stage the file under $ONNX_DIR rather than modifying the Hugging Face cache; the builder carries it into the engine bundle together with policy.json. Fetching it into $ENGINE_DIR instead does not work: an engine directory holding only an untagged assets/ is refused as a build target.

2. Build all engines#

"$BUILD_DIR/experimental_models/pi05/examples/pi05_policy_build" \
    --onnxDir "$ONNX_DIR" --engineDir "$ENGINE_DIR"

$ENGINE_DIR must be empty or already hold this same export; building a different export into a populated directory is refused, so use a fresh one. Rebuilding components from the export already staged there keeps working. --component visual|prefix|action|cond builds a single component, and --maxBatchSize N widens the batch axis.

3. Run the policy#

"$BUILD_DIR/experimental_models/pi05/examples/pi05_policy_inference" \
    --engineDir "$ENGINE_DIR" \
    --inputFile observation.json \
    --output action.json
{
  "task": "pick up the black bowl",
  "state": [0.02, -0.11, 0.98, 3.05, -0.21, 0.04, 0.031, -0.029],
  "cameras": {
    "observation/image": "base.png",
    "observation/wrist_image": "wrist.png"
  }
}

DROID names its two cameras observation/exterior_image_1_left and observation/wrist_image_left, and may spell its state as joint_position (7) plus gripper_position (1) instead of the joined state; carrying both spellings is rejected. ALOHA names cam_high, cam_left_wrist and cam_right_wrist, of which only the first is required, and accepts cam_low while dropping it, as openpi does.

{
  "task": "fold the towel",
  "state": [0.998, -0.07, -0.793, -0.468, 0.523, -0.726, 0.35,
            -0.504, -0.898, 0.135, -0.901, -0.173, -0.447, 0.72],
  "cameras": {"cam_high": "high.png", "cam_left_wrist": "left.png"}
}

This request shape is the example CLI’s own, not openpi’s: openpi’s policies are served prompt alongside observation/-prefixed state and image keys. What matches openpi is the policy behind it – the prompt, the normalization and each configuration’s conversions – so a client written against openpi’s server maps its fields onto these rather than posting them unchanged.

cameras is keyed by the names policy.json declares, so an unknown or repeated name is rejected; state is in the robot’s native units. Only the cameras a request supplies are run – an optional slot left out is not processed at all. --steps sets the denoise step count, --cudaGraph captures the loop, and --batch B replicates the request across engines built with --maxBatchSize B.

action.json holds actions, the normalized zero-padded [action_horizon, action_dim] chunk, and robot_actions, the same chunk in robot units sliced to the embodiment width. A canonical-tensor run writes actions alone: converting to robot units needs the request’s own state, which ALOHA’s absolute-action step reads and those tensors do not carry.

4. Verify correctness#

This is a numerical comparison against openpi’s PyTorch chunk for one fixed observation, not task success. Run both sides on the same preprocessed pixels, token ids and initial noise, at the same denoise-step count (--steps here, num_inference_steps there), then compare the normalized chunks. x_0 must come from the reference run: the two implementations draw from different RNGs, so a shared --seed does not give a shared trajectory.

# Same tensors both sides saw, so any difference is the engines.
"$BUILD_DIR/experimental_models/pi05/examples/pi05_policy_inference" \
    --engineDir "$ENGINE_DIR" \
    --pixelValues ref/pixel_values.bin --tokens "$(cat ref/token_ids.csv)" \
    --noise ref/x0.bin --output actual.safetensors

python experimental_models/pi05/scripts/compare_pi05_actions.py \
    --reference ref/actions.npy --actual actual.safetensors

Drop --pixelValues/--tokens for --inputFile observation.json to put the observation front end back in the loop and score that instead.

The comparator reports cosine similarity and max absolute error over the normalized chunk and exits non-zero outside its thresholds; --json-field robot_actions scores a JSON --output in robot units instead, which needs no openpi internals. pi0.5 Design covers where each reference tensor comes from and what the comparison does and does not prove.