Artificial intelligence
Flow matching
Definition
Flow matching is a generative-model training method that learns a vector field for transforming a simple probability distribution into a data distribution. In robot learning, the generated samples can be continuous action sequences conditioned on observations and instructions.
Also known as: FM
Updated
Learn how a sample should move
Flow Matching for Generative Modeling trains a model by regressing vector fields along chosen probability paths between noise and data. At generation time, a numerical solver follows the learned field to transform an initial sample into a data-like sample.
Here, flow refers to movement through a mathematical sample space. It does not mean fluid flow around a robot or optical flow between camera frames.
Continuous robot actions are one application
The pi0 model combines a pretrained vision-language model with a flow-matching action architecture. Visual and language information condition the generation of robot actions. Its authors evaluate manipulation tasks including laundry folding and box assembly.
An action sample can represent several future commands, making this approach compatible with action chunking.
Flow matching and diffusion overlap
The original flow-matching formulation can use diffusion probability paths as well as other paths. The terms therefore describe related families of generative methods rather than completely separate ideas.
A chosen path, solver, and number of integration steps affect generation. Results from one model or solver configuration should not be treated as evidence that every flow-matching policy is faster or more accurate than every diffusion policy.
Sources
Related terms
Diffusion policy
A diffusion policy generates robot actions through a learned denoising process conditioned on observations. It commonly predicts an action sequence by progressively refining a noisy candidate rather than predicting one action with a single direct regression.
Action chunking
Action chunking is the prediction or organization of several future robot actions as one sequence. A policy can execute all or part of a chunk before using new observations to produce another sequence.
Vision-language-action model
A vision-language-action model is an AI model that uses visual observations and language instructions to produce actions for a robot. It connects what a robot sees and what it is asked to do with outputs that a robot controller can execute.