Vision-Language-Action#
Vision-language-action models use model-specific action contracts rather than the standard text-generation output. Choose the guide for the checkpoint and runtime:
Model |
Input |
Output |
Executable |
|---|---|---|---|
camera frames, instruction, past trajectory |
future acceleration/curvature trajectory |
|
|
observation image or frame list, instruction |
robot action chunk |
|
Both workflows export on CPU, build all required TensorRT engines on the target, and invoke one end-to-end runtime executable.