pi0.5#
pi0.5 is Physical Intelligence’s flow-matching Vision-Language-Action model. Edge-LLM runs it as four TensorRT components driven by an experimental C++ runtime, which returns an action chunk for a set of camera views and an instruction. See pi0.5 Design for the component contracts and the reasoning behind the split.
This workflow supports three openpi configurations, in FP16, and follows openpi’s policy
contract for each. The Hugging Face mirror is a weight distribution channel – its own LeRobot
policy configuration is a different contract and is not used. Results are a numerical comparison
against openpi’s action chunk, and all three configurations meet the comparator’s gate on their own
embodiment’s observations. Task success is not claimed. The exporter consumes an already-converted PyTorch
checkpoint (model.safetensors plus config.json); the LeRobot releases are in that form.
|
weights |
norm stats |
state / action |
horizon |
cameras |
|---|---|---|---|---|---|
|
|
|
8 / 7 |
10 |
two, both required |
|
|
|
8 / 8 |
15 |
exterior and left wrist, both required |
|
|
|
14 / 14 |
50 |
|
openpi publishes no pi05_aloha checkpoint: its configuration is an inference contract over the
generalist pi05_base weights, which is what the row above names.
Complete Installation first; pi0.5 also requires
-DBUILD_EXPERIMENTAL_MODELS=ON. Its action engine carries a trt_edgellm::AttentionPlugin node, so
EDGELLM_PLUGIN_PATH must name the built plugin library or that engine cannot be built or
deserialized.
export CHECKPOINT=lerobot/pi05_libero_base
export POLICY_CONFIG=pi05_libero
export ONNX_DIR=$HOME/pi05_onnx
export ENGINE_DIR=$HOME/pi05_engines
export BUILD_DIR=/path/to/tensorrt-edge-llm/build
export EDGELLM_PLUGIN_PATH="$BUILD_DIR/libNvInfer_edgellm_plugin.so"
1. Export on CPU#
export PYTHONNOUSERSITE=1
tensorrt-edgellm-export "$CHECKPOINT" "$ONNX_DIR" --dtype float16 \
--pi05-policy-config "$POLICY_CONFIG"
--pi05-policy-config fixes the action horizon, the camera slots and the prompt form. A converted
checkpoint names no configuration, so the flag is required for every one but pi05_libero, whose
feature contract names itself. The checkpoint’s own n_action_steps is cross-checked against the
configuration, and a pairing whose horizons disagree is refused before anything is built.
This writes one subdirectory per component – visual/, prefix/, action/ and cond/ – plus
text_tokenizer/, the token-embedding table, and policy.json, the observation and action
contract: shapes from the checkpoint, policy semantics from the openpi configuration. Both the
builder and the runtime refuse a bundle without policy.json. The tokenizer is fetched from
gs://big_vision/paligemma_tokenizer.model, which openpi loads the same way; the Hugging Face
mirror is gated and cannot be fetched unattended.
cond/ precomputes the AdaRMS modulation schedule instead of rebuilding it inside the per-step
action graph. The runtime evaluates it only when the denoise-step count or the batch changes, so a
run at fixed settings pays it once, not once per request. It is optional:
--no-pi05-hoist-adarms-cond exports the other three alone, for debugging and A/B runs.
The LeRobot checkpoints are PyTorch conversions of openpi’s. They carry the processor schemas but not their normalization state – those statistics are policy input/output metadata, not TensorRT calibration data, and they stay in the original openpi checkpoint assets. The runtime refuses to load without them, so fetch the file for the configuration being exported, from the table above:
mkdir -p "$ONNX_DIR/assets"
curl --fail --location --output "$ONNX_DIR/assets/norm_stats.json" \
https://storage.googleapis.com/openpi-assets/checkpoints/pi05_libero/assets/physical-intelligence/libero/norm_stats.json
# pi05_droid:
# https://storage.googleapis.com/openpi-assets/checkpoints/pi05_droid/assets/droid/norm_stats.json
# pi05_aloha:
# https://storage.googleapis.com/openpi-assets/checkpoints/pi05_base/assets/trossen/norm_stats.json
Stage the file under $ONNX_DIR rather than modifying the Hugging Face cache; the builder carries it
into the engine bundle together with policy.json. Fetching it into $ENGINE_DIR instead does not
work: an engine directory holding only an untagged assets/ is refused as a build target.
2. Build all engines#
"$BUILD_DIR/experimental_models/pi05/examples/pi05_policy_build" \
--onnxDir "$ONNX_DIR" --engineDir "$ENGINE_DIR"
$ENGINE_DIR must be empty or already hold this same export; building a different export into a
populated directory is refused, so use a fresh one. Rebuilding components from the export already
staged there keeps working. --component visual|prefix|action|cond builds a single component, and
--maxBatchSize N widens the batch axis.
3. Run the policy#
"$BUILD_DIR/experimental_models/pi05/examples/pi05_policy_inference" \
--engineDir "$ENGINE_DIR" \
--inputFile observation.json \
--output action.json
{
"task": "pick up the black bowl",
"state": [0.02, -0.11, 0.98, 3.05, -0.21, 0.04, 0.031, -0.029],
"cameras": {
"observation/image": "base.png",
"observation/wrist_image": "wrist.png"
}
}
DROID names its two cameras observation/exterior_image_1_left and observation/wrist_image_left,
and may spell its state as joint_position (7) plus gripper_position (1) instead of the joined
state; carrying both spellings is rejected. ALOHA names cam_high, cam_left_wrist and
cam_right_wrist, of which only the first is required, and accepts cam_low while dropping it, as
openpi does.
{
"task": "fold the towel",
"state": [0.998, -0.07, -0.793, -0.468, 0.523, -0.726, 0.35,
-0.504, -0.898, 0.135, -0.901, -0.173, -0.447, 0.72],
"cameras": {"cam_high": "high.png", "cam_left_wrist": "left.png"}
}
This request shape is the example CLI’s own, not openpi’s: openpi’s policies are served
prompt alongside observation/-prefixed state and image keys. What matches openpi is the policy
behind it – the prompt, the normalization and each configuration’s conversions – so a client
written against openpi’s server maps its fields onto these rather than posting them unchanged.
cameras is keyed by the names policy.json declares, so an unknown or repeated name is rejected;
state is in the robot’s native units. Only the cameras a request supplies are run – an
optional slot left out is not processed at all. --steps sets the denoise step count, --cudaGraph
captures the loop, and --batch B replicates the request across engines built with
--maxBatchSize B.
action.json holds actions, the normalized zero-padded [action_horizon, action_dim] chunk, and
robot_actions, the same chunk in robot units sliced to the embodiment width. A canonical-tensor run
writes actions alone: converting to robot units needs the request’s own state, which ALOHA’s
absolute-action step reads and those tensors do not carry.
4. Verify correctness#
This is a numerical comparison against openpi’s PyTorch chunk for one fixed observation, not task
success.
Run both sides on the same preprocessed pixels, token ids and initial noise, at the same denoise-step
count (--steps here, num_inference_steps there), then compare the normalized chunks. x_0 must
come from the reference run: the two implementations draw from different RNGs, so a shared --seed
does not give a shared trajectory.
# Same tensors both sides saw, so any difference is the engines.
"$BUILD_DIR/experimental_models/pi05/examples/pi05_policy_inference" \
--engineDir "$ENGINE_DIR" \
--pixelValues ref/pixel_values.bin --tokens "$(cat ref/token_ids.csv)" \
--noise ref/x0.bin --output actual.safetensors
python experimental_models/pi05/scripts/compare_pi05_actions.py \
--reference ref/actions.npy --actual actual.safetensors
Drop --pixelValues/--tokens for --inputFile observation.json to put the observation front end
back in the loop and score that instead.
The comparator reports cosine similarity and max absolute error over the normalized chunk and exits
non-zero outside its thresholds; --json-field robot_actions scores a JSON --output in robot units
instead, which needs no openpi internals.
pi0.5 Design covers where each reference tensor comes
from and what the comparison does and does not prove.