Artificial intelligence
Vision-language-action model
Definition
A vision-language-action model is an AI model that uses visual observations and language instructions to produce actions for a robot. It connects what a robot sees and what it is asked to do with outputs that a robot controller can execute.
Also known as: VLA, VLA model, Vision language action model
Updated
Action is part of the model output
A vision-language model can describe an image or answer a question about it. A VLA also produces robot action outputs. RT-2 represents actions as tokens, allowing a model trained with visual and language data to generate instructions for robotic control.
For example, a robot may receive a camera view and an instruction to move an object. The output needs to guide movement, not merely describe the object.
Different models can use different robot bodies
Gemini Robotics is described by its developers as a VLA model. Their work includes two-arm platforms and adaptation to other robot forms. A VLA is therefore a model type, not a synonym for a humanoid robot.
A model is part of a robot system
Camera input, action representation, controllers, hardware, and operating conditions affect what a system can do. The RT-2 and Gemini Robotics reports describe particular training setups and evaluations. Their results should be read in that context, rather than as evidence that every VLA can perform every physical task.
Sources
Related terms
Embodied AI
Embodied AI is artificial intelligence that perceives and acts through a body in a physical or simulated environment. It connects sensing, reasoning, and action rather than producing only text or images.
Robotic manipulation
Robotic manipulation is the use of a robot to change an object's position, orientation, or state through physical interaction. It includes grasping and moving objects as well as actions such as pushing or carrying them without a grasp.
Humanoid robot
A humanoid robot is a robot with a body arranged to resemble the human form, usually with a torso, arms, and legs. The term describes its physical form and does not by itself establish human-level intelligence or general autonomy.