Artificial intelligence
Dataset aggregation
Definition
Dataset aggregation, usually called DAgger in imitation learning, is an iterative algorithm that collects expert action labels at states visited by a learner. It adds those examples to an accumulated dataset and retrains the policy.
Also known as: DAgger
Updated
Train on states the learner encounters
The DAgger algorithm alternates between executing a policy, obtaining the expert's preferred actions for visited states, and training on the combined dataset. Its name abbreviates Dataset Aggregation.
The original formulation can mix expert and learner behavior during data collection. The expert supplies action labels even when the learner has caused the system to enter a state outside the original demonstrations.
Correct the distribution mismatch
Basic behavior cloning learns from states produced by expert behavior. Its own errors can later move it into different states. DAgger explicitly gathers training examples from the distributions induced during its iterative learning process.
For a reaching robot, this could mean asking for a corrective action after the learned controller approaches an object from an awkward pose. Adding another perfectly executed demonstration might never show that pose.
Expert access is a real requirement
DAgger is more specific than simply combining robot datasets. Its defining feature is the loop between learner execution, expert labeling, aggregation, and retraining.
The original analysis gives performance guarantees under stated learning and reduction assumptions. It does not guarantee safe exploration on physical hardware. Deploying the collection loop requires an expert who can label encountered states and a suitable way to manage the learner's actions during those trials.
Sources
Related terms
Imitation learning
Imitation learning learns behavior from examples supplied by a demonstrator. In robotics, demonstrations can teach a policy how to perform a task without requiring every action or objective to be programmed by hand.
Behavior cloning
Behavior cloning is an imitation-learning method that trains a policy to predict a demonstrator’s actions from recorded observations or states. It treats action prediction as a supervised-learning problem.
Teleoperation
Teleoperation is the control of a robot by a human operator from a separate location or interface. The operator supplies commands while feedback, such as camera images or the robot's motion, helps them guide the task.