Vision-Language-Action Model (VLA)
A vision-language-action model is a neural network that takes camera images and a natural-language instruction as input and outputs robot actions, typically built by extending a vision-language model and training it on robot demonstration data.
VLAs start from a vision-language model pre-trained on large image and text datasets and are further trained on robot trajectories that pair observations and instructions with actions. Actions may be represented as discrete tokens, as in Google DeepMind's RT-2, which introduced the term, or produced by continuous action decoders such as diffusion or flow-matching heads. The model runs in a closed loop, predicting the next action or a short chunk of actions from the latest observation.
The goal is generalist robot policies that follow open-ended instructions, apply knowledge about objects learned from web-scale data to manipulation and adapt to new tasks with little additional data. Examples include RT-2 and the open-source OpenVLA, and VLAs are being developed for robot arms, mobile manipulators and humanoid robots.
VLAs are large models that need capable accelerators for low-latency inference, and their success rates vary widely across tasks and environments, so reliability must be measured carefully. Data covering the target robot and setting is usually required. Because their behaviour cannot be fully verified, they must run within safety-rated limits, with safety functions independent of the model.
Key points
- Maps images and language instructions to robot actions
- Built on vision-language models trained further on robot demonstrations
- RT-2 introduced the term; OpenVLA is an open-source example
- Safety functions must remain independent of the model
Where AiVibe comes in
AiVibe designs and manufactures the AiAmbA AI Factory, whose edge devices and AI agents let people talk to robot controllers in plain language. Robotics perception is an AiAmbA use case, and robot safety functions follow ISO 10218 and never depend on the AI layer.
Related terms
- Physical AIRobotics & Physical AI
- Imitation LearningRobotics & Physical AI
- Humanoid RobotRobotics & Physical AI
- Large Language Model (LLM)AI & Machine Learning
- Computer VisionAI & Machine Learning
- AI AgentAI & Machine Learning