Artificial intelligence
Generalist robot policy
Definition
A generalist robot policy is a learned action-selection model designed to perform multiple tasks across a range of robot settings. Its generality depends on the tasks, observations, action interfaces, and robot bodies included in training and evaluation.
Also known as: Generalist robotic policy
Updated
One policy covers a range of tasks
A policy maps observations and task information to actions. A generalist policy shares this mapping across tasks instead of assigning a separately trained model to every task. Octo accepts language instructions or goal images and uses a shared model for multiple manipulation settings.
For a robot arm, the same model might receive different goals for picking up an object, moving it to a container, or inserting a part. The instruction or goal distinguishes the requested behavior.
Generality has several dimensions
Task variety, unfamiliar objects, changed cameras, and a different robot body are separate challenges. Octo's authors distinguish direct evaluation in training-related setups from fine-tuning to new observations and action spaces. Success on a new task with the same arm does not by itself demonstrate transfer to a humanoid hand.
Relationship to foundation models
The pi0 paper uses generalist policy and robot foundation model as overlapping terms. A useful distinction is that generalist describes a policy's intended scope, while foundation describes its role as a reusable pretrained model.
Read reported results together with the supported action representation, adaptation data, and evaluated tasks. The label alone does not specify how many tasks the policy can reliably perform.
Sources
Related terms
Robot foundation model
A robot foundation model is a model pretrained on broad data to support adaptation to multiple robot tasks, environments, or bodies. The term describes a reusable learning base rather than a guarantee of general physical competence.
Language-conditioned policy
A language-conditioned policy selects actions using a language instruction together with observations. The instruction specifies or modifies the behavior requested from the policy.
Cross-embodiment learning
Cross-embodiment learning uses experience from different robot bodies to train representations or policies that can transfer across those bodies. It requires a way to handle differences in sensing, geometry, and available actions.