Artificial intelligence
Diffusion policy
Definition
A diffusion policy generates robot actions through a learned denoising process conditioned on observations. It commonly predicts an action sequence by progressively refining a noisy candidate rather than predicting one action with a single direct regression.
Updated
Denoising produces an action sequence
The Diffusion Policy research models visuomotor behavior as a conditional denoising diffusion process. During training, the model learns to remove noise from demonstrated actions using observations as context. During execution, repeated refinement turns an initial noisy sequence into a candidate robot motion.
The noisy sequence exists inside the model's computation. It is not a command to make the physical robot move randomly.
Multiple valid motions can remain distinct
A robot may be able to push an object around either side of an obstacle. A model trained to average incompatible demonstrations could propose an unsuitable middle path. Diffusion Policy is designed to represent multiple modes of an action distribution and select a coherent sequence, as illustrated in the authors' manipulation experiments.
It also combines action chunking with receding-horizon execution: the controller executes part of a predicted sequence, observes again, and generates another sequence.
Sampling and feedback impose practical limits
Iterative generation takes computation, so sampling settings and observation timing affect the control loop. The paper's demonstrations include pushing, mug flipping, and sauce manipulation. These results support those evaluated setups; diffusion decoding alone does not provide a collision guarantee or establish humanoid walking capability.
Sources
Related terms
Action chunking
Action chunking is the prediction or organization of several future robot actions as one sequence. A policy can execute all or part of a chunk before using new observations to produce another sequence.
Behavior cloning
Behavior cloning is an imitation-learning method that trains a policy to predict a demonstrator’s actions from recorded observations or states. It treats action prediction as a supervised-learning problem.
Flow matching
Flow matching is a generative-model training method that learns a vector field for transforming a simple probability distribution into a data distribution. In robot learning, the generated samples can be continuous action sequences conditioned on observations and instructions.