Artificial intelligence

Behavior cloning

Definition

Behavior cloning is an imitation-learning method that trains a policy to predict a demonstrator’s actions from recorded observations or states. It treats action prediction as a supervised-learning problem.

Also known as: Behaviour cloning, Behavioral cloning, Behavioural cloning, BC

Updated

Fit actions to demonstrated observations

A training example pairs what the demonstrator observed with the action taken. The learner adjusts its predictions to match those examples. The DAgger paper's supervised-imitation formulation formalizes this approach under the distribution of states visited by the expert.

For a robot arm, the input might contain camera images and joint positions, while the target is a joint command or an end-effector movement. Predicting an action sequence rather than one action remains compatible with supervised imitation, as illustrated by ACT.

Small mistakes can change later inputs

Robot actions influence what the policy observes next. If a learned grasp misses slightly, the next image may differ from anything in the demonstrations. Further errors can then accumulate.

The DAgger analysis explains why low prediction error on expert data does not automatically imply low error over an entire executed task.

Distinguish copying from reward optimization

Behavior cloning learns to match demonstrated actions. Offline reinforcement learning instead uses previously collected experience to optimize a reward-based objective. Both can use a fixed dataset, but their objectives differ.

Dataset aggregation addresses a specific weakness of basic cloning by collecting expert labels for states the learner actually visits. It requires additional interaction and expert access.

Sources