Artificial intelligence
Policy distillation
Definition
Policy distillation trains a student policy to reproduce behavior from one or more teacher policies. It can transfer learned behavior into a smaller network or combine multiple task-specific policies into one model.
Updated
Transfer behavior from teacher to student
The Policy Distillation paper presents distillation as a way to extract a reinforcement-learning agent's policy into a new network. The training signal comes from the teacher's learned behavior rather than requiring the student to rediscover that behavior through the original training process.
In a robotics application, an expensive teacher policy could guide training of a smaller policy intended for a robot's onboard computer. That is a possible application of the method, not a hardware result established by the original paper.
Compression and consolidation are different aims
A student may be smaller than one teacher. Alternatively, it may combine knowledge from several task-specific teachers into a shared policy. The original work studies both aims using Atari tasks.
Combining teachers is relevant to a generalist robot policy, but distillation alone does not establish that the combined model resolves conflicts between tasks or handles unfamiliar situations.
Evaluate the student independently
The student approximates its teachers through its own architecture and training data. Their abilities do not transfer automatically or exactly.
As with behavior cloning, matching actions on sampled inputs is only part of the evaluation. A robot policy must also be assessed while acting, when its own decisions influence later observations. The original paper's compression results should not be assumed to hold for a different robot or student model.
Sources
Related terms
Reinforcement learning
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
Behavior cloning
Behavior cloning is an imitation-learning method that trains a policy to predict a demonstrator’s actions from recorded observations or states. It treats action prediction as a supervised-learning problem.
Generalist robot policy
A generalist robot policy is a learned action-selection model designed to perform multiple tasks across a range of robot settings. Its generality depends on the tasks, observations, action interfaces, and robot bodies included in training and evaluation.