Artificial intelligence

Policy distillation

Definition

Policy distillation trains a student policy to reproduce behavior from one or more teacher policies. It can transfer learned behavior into a smaller network or combine multiple task-specific policies into one model.

Updated

Transfer behavior from teacher to student

The Policy Distillation paper presents distillation as a way to extract a reinforcement-learning agent's policy into a new network. The training signal comes from the teacher's learned behavior rather than requiring the student to rediscover that behavior through the original training process.

In a robotics application, an expensive teacher policy could guide training of a smaller policy intended for a robot's onboard computer. That is a possible application of the method, not a hardware result established by the original paper.

Compression and consolidation are different aims

A student may be smaller than one teacher. Alternatively, it may combine knowledge from several task-specific teachers into a shared policy. The original work studies both aims using Atari tasks.

Combining teachers is relevant to a generalist robot policy, but distillation alone does not establish that the combined model resolves conflicts between tasks or handles unfamiliar situations.

Evaluate the student independently

The student approximates its teachers through its own architecture and training data. Their abilities do not transfer automatically or exactly.

As with behavior cloning, matching actions on sampled inputs is only part of the evaluation. A robot policy must also be assessed while acting, when its own decisions influence later observations. The original paper's compression results should not be assumed to hold for a different robot or student model.

Sources