Artificial intelligence
Robot foundation model
Definition
A robot foundation model is a model pretrained on broad data to support adaptation to multiple robot tasks, environments, or bodies. The term describes a reusable learning base rather than a guarantee of general physical competence.
Also known as: Robotics foundation model
Updated
Pretraining creates a reusable starting point
The foundation-model report defines foundation models through broad training data and adaptation to many downstream tasks. In robotics, this can mean reusing learned visual features, action patterns, or both when training a robot for a new setting.
For example, a pretrained manipulation model may provide the starting weights for learning to sort objects on a different robot arm. The new robot can still need its own demonstrations, sensor configuration, and action interface.
Related labels describe different properties
Some robotics papers use this label interchangeably with generalist robot policy, including the pi0 paper. Foundation emphasizes reuse through pretraining; generalist emphasizes the range of tasks a policy performs. Neither label specifies one architecture.
A vision-language-action model instead describes the relationship between visual input, language, and action output. These descriptions can apply to the same model.
Evaluate the adaptation that was demonstrated
Octo reports both control in settings represented in its training data and fine-tuning to new setups. Those are different tests. A broad training mixture does not establish that a model can operate an unfamiliar humanoid without adaptation or task-specific evaluation.
Sources
Related terms
Generalist robot policy
A generalist robot policy is a learned action-selection model designed to perform multiple tasks across a range of robot settings. Its generality depends on the tasks, observations, action interfaces, and robot bodies included in training and evaluation.
Vision-language-action model
A vision-language-action model is an AI model that uses visual observations and language instructions to produce actions for a robot. It connects what a robot sees and what it is asked to do with outputs that a robot controller can execute.
Cross-embodiment learning
Cross-embodiment learning uses experience from different robot bodies to train representations or policies that can transfer across those bodies. It requires a way to handle differences in sensing, geometry, and available actions.