Artificial intelligence
Imitation learning
Definition
Imitation learning learns behavior from examples supplied by a demonstrator. In robotics, demonstrations can teach a policy how to perform a task without requiring every action or objective to be programmed by hand.
Also known as: Learning from demonstration, Learning from demonstrations, LfD
Updated
Demonstrations provide the learning signal
An Algorithmic Perspective on Imitation Learning describes learning from demonstrations as an alternative to manually engineering complex behavior. Demonstrations provide evidence about how a task should be performed, although different algorithms use that evidence differently.
For example, a person can use teleoperation to guide two robot arms through an insertion task. The ALOHA and ACT project records real demonstrations and trains a policy to reproduce related behaviors.
Behavior cloning is one method
Behavior cloning directly trains action predictions from demonstrated observations and actions. Imitation learning is the broader field; it also includes interactive approaches and methods that use demonstrations to define a learning objective.
A motion-imitation system can even use reinforcement learning to follow a reference trajectory, as in Peng and colleagues' locomotion work. Imitation and reinforcement learning therefore need not be mutually exclusive.
Reproduction depends on data and embodiment
The demonstrator's motions, the robot's observations, and the robot's possible actions must be connected. Human motion may require retargeting, while robot demonstrations can still omit recovery from errors.
A successful replay or training example does not establish general task competence. Evaluation should test the learned policy under the conditions in which it will actually act.
Sources
Related terms
Behavior cloning
Behavior cloning is an imitation-learning method that trains a policy to predict a demonstrator’s actions from recorded observations or states. It treats action prediction as a supervised-learning problem.
Dataset aggregation
Dataset aggregation, usually called DAgger in imitation learning, is an iterative algorithm that collects expert action labels at states visited by a learner. It adds those examples to an accumulated dataset and retrains the policy.
Teleoperation
Teleoperation is the control of a robot by a human operator from a separate location or interface. The operator supplies commands while feedback, such as camera images or the robot's motion, helps them guide the task.