AiVibe

Robotics & Physical AI

Vision-Language-Action Model (VLA)

A vision-language-action model is a neural network that takes camera images and a natural-language instruction as input and outputs robot actions, typically built by extending a vision-language model and training it on robot demonstration data.

VLAs start from a vision-language model pre-trained on large image and text datasets and are further trained on robot trajectories that pair observations and instructions with actions. Actions may be represented as discrete tokens, as in Google DeepMind's RT-2, which introduced the term, or produced by continuous action decoders such as diffusion or flow-matching heads. The model runs in a closed loop, predicting the next action or a short chunk of actions from the latest observation.

The goal is generalist robot policies that follow open-ended instructions, apply knowledge about objects learned from web-scale data to manipulation and adapt to new tasks with little additional data. Examples include RT-2 and the open-source OpenVLA, and VLAs are being developed for robot arms, mobile manipulators and humanoid robots.

VLAs are large models that need capable accelerators for low-latency inference, and their success rates vary widely across tasks and environments, so reliability must be measured carefully. Data covering the target robot and setting is usually required. Because their behaviour cannot be fully verified, they must run within safety-rated limits, with safety functions independent of the model.

Key points

Where AiVibe comes in

AiVibe designs and manufactures the AiAmbA AI Factory, whose edge devices and AI agents let people talk to robot controllers in plain language. Robotics perception is an AiAmbA use case, and robot safety functions follow ISO 10218 and never depend on the AI layer.

Explore AiVibe’s work in Robotics & Physical AI →

Related terms

Terms that refer to Vision-Language-Action Model (VLA)

Ask AiMuruga can explain Vision-Language-Action Model (VLA) for your plant, product or security programme, and draw how it fits.