Artificial intelligence
Behavior cloning
Definition
Behavior cloning is an imitation-learning method that trains a policy to predict a demonstrator’s actions from recorded observations or states. It treats action prediction as a supervised-learning problem.
Also known as: Behaviour cloning, Behavioral cloning, Behavioural cloning, BC
Updated
Fit actions to demonstrated observations
A training example pairs what the demonstrator observed with the action taken. The learner adjusts its predictions to match those examples. The DAgger paper's supervised-imitation formulation formalizes this approach under the distribution of states visited by the expert.
For a robot arm, the input might contain camera images and joint positions, while the target is a joint command or an end-effector movement. Predicting an action sequence rather than one action remains compatible with supervised imitation, as illustrated by ACT.
Small mistakes can change later inputs
Robot actions influence what the policy observes next. If a learned grasp misses slightly, the next image may differ from anything in the demonstrations. Further errors can then accumulate.
The DAgger analysis explains why low prediction error on expert data does not automatically imply low error over an entire executed task.
Distinguish copying from reward optimization
Behavior cloning learns to match demonstrated actions. Offline reinforcement learning instead uses previously collected experience to optimize a reward-based objective. Both can use a fixed dataset, but their objectives differ.
Dataset aggregation addresses a specific weakness of basic cloning by collecting expert labels for states the learner actually visits. It requires additional interaction and expert access.
Sources
Related terms
Imitation learning
Imitation learning learns behavior from examples supplied by a demonstrator. In robotics, demonstrations can teach a policy how to perform a task without requiring every action or objective to be programmed by hand.
Dataset aggregation
Dataset aggregation, usually called DAgger in imitation learning, is an iterative algorithm that collects expert action labels at states visited by a learner. It adds those examples to an accumulated dataset and retrains the policy.
Offline reinforcement learning
Offline reinforcement learning learns a reward-optimizing policy from previously collected experience without gathering new environment interactions during that learning stage. The data may come from earlier policies, demonstrations, or other collection procedures.