Skip to main content

VLA Vision-Language-Action Model

VLA (Vision-Language-Action) models are end-to-end policy models that combine visual perception, natural-language instructions, and robot action control. Given camera images, robot proprioceptive state, and a natural-language task description, the model directly outputs end-effector and gripper control actions for imitation learning and real-robot deployment.

This section is based on the OpenPI framework and describes the full workflow from data collection, dataset conversion, and model training to real-robot inference.

Applicable Scenarios​

  • Fine-tune PI0 / PI05 models using LeRobot-format datasets
  • Convert LeRobot v3.0 datasets exported by Data-Processing-Tool to the v2.1 format required by OpenPI training
  • Deploy trained checkpoints on robots such as Marvin Pro and execute language-instruction tasks through a WebSocket policy service

End-to-End Flow​

Raw acquisition data (BAG / MCAP)
-> Data-Processing-Tool exports LeRobot v3.0 dataset
-> convert_v3_to_v2.py converts to v2.1 format
-> train_pytorch.py trains PI0 / PI05 model
-> serve_policy_kmd_joint.py starts the policy server
-> vlahost collects robot state and camera images
-> openpi_client_policy_http_kmd_joint.py performs inference and sends actions

Observation and Action Format​

The current KMD real-robot deployment uses three cameras and a 16D joint-space state:

FieldTypeDescription
statefloat32 vector16D state: 7 left-arm joints, left gripper, 7 right-arm joints, right gripper
cam_highHWC uint8Main-view camera image
cam_left_wristHWC uint8Left-wrist camera image
cam_right_wristHWC uint8Right-wrist camera image
promptstringNatural-language task instruction

Policy output:

FieldTypeDescription
actions[T, 16]Joint-action chunk in the same order as the 16D state

The model uses degrees internally. Robot-side joint_states.positions normally uses radians. The OpenPI client converts rad -> deg before inference and deg -> rad before sending an action.

Environment Requirements​

  • Python 3.10+
  • uv package manager
  • Run all commands from the OpenPI project root
  • NVIDIA GPU is recommended for training; inference can run through a policy service on an independent GPU server

Quick Navigation​

DocumentContent
Dataset ConversionLeRobot v3.0 to v2.1 format conversion
Dataset SampleFour-camera video and directory structure description
Environment and Project SetupOpenPI KMD, GPU, Docker, and robot-side vlahost environment
Model TrainingPI0 / PI05 fine-tuning and checkpoint management
Real-Robot DeploymentPolicy server and robot client integration