Artificial intelligence

Vision-language-action model

Definition

A vision-language-action model is an AI model that uses visual observations and language instructions to produce actions for a robot. It connects what a robot sees and what it is asked to do with outputs that a robot controller can execute.

Also known as: VLA, VLA model, Vision language action model

Updated

Action is part of the model output

A vision-language model can describe an image or answer a question about it. A VLA also produces robot action outputs. RT-2 represents actions as tokens, allowing a model trained with visual and language data to generate instructions for robotic control.

For example, a robot may receive a camera view and an instruction to move an object. The output needs to guide movement, not merely describe the object.

Different models can use different robot bodies

Gemini Robotics is described by its developers as a VLA model. Their work includes two-arm platforms and adaptation to other robot forms. A VLA is therefore a model type, not a synonym for a humanoid robot.

A model is part of a robot system

Camera input, action representation, controllers, hardware, and operating conditions affect what a system can do. The RT-2 and Gemini Robotics reports describe particular training setups and evaluations. Their results should be read in that context, rather than as evidence that every VLA can perform every physical task.

Sources