Vision-Language-Action#
Vision-language-action models use model-specific action contracts rather than the standard text-generation output. Choose the guide for the checkpoint and runtime:
Model |
Input |
Output |
Executable |
|---|---|---|---|
camera frames, instruction, past trajectory |
future acceleration/curvature trajectory |
|
|
observation image or frame list, instruction |
robot action chunk |
|
|
camera views, instruction |
robot action chunk |
|
Each workflow exports on CPU, builds all required TensorRT engines on the target, and invokes one end-to-end runtime executable.