Artificial intelligence

Privileged information

Definition

Privileged information is training-only information that a robot policy will not receive when it is deployed. Simulator state, exact terrain geometry, or object pose can help train a critic, teacher, or auxiliary model while the deployed policy uses available sensors.

Also known as: Learning with privileged information, Privileged observations

Updated

Information available only during training

A simulator exposes quantities that a physical robot may not measure directly, including exact contact states, terrain heights, object poses, and complete system state. These signals are privileged when an algorithm can use them in training but the deployed policy cannot depend on them.

In asymmetric actor-critic training, the critic receives the simulator's full state while the actor receives rendered RGB-D observations. The actor is therefore trained for the observation interface available at deployment, while the critic uses extra state to estimate training targets. The authors reported simulation-to-real experiments for picking, pushing, and moving a block, but that result belongs to their particular tasks and training setup.

Teachers and students can see different inputs

Privileged information can also be given to an expert or teacher. A student then learns to reproduce useful behaviour from deployable inputs such as cameras and proprioception. This is one use of policy distillation, but the two concepts are not identical: distillation describes transferring behaviour, while privilege describes an information asymmetry during training.

The 2026 iGPC paper describes interaction experts conditioned on privileged scene state and a student driven by onboard sensory observations. Its reported humanoid interaction experiments are evidence for that framework, not proof that any choice of privileged state will improve a policy.

Deployment inputs remain the real constraint

Training-only signals can make optimisation easier without making those signals observable on hardware. If the student cannot infer a necessary distinction from its sensors or history, the teacher's additional knowledge cannot remove that ambiguity. A policy may also learn correlations with simulated privileged state that fail under real dynamics.

Evaluation should confirm that privileged inputs are disconnected at deployment and should test sensor noise, delay, occlusion, and model mismatch. Privileged training can support sim-to-real transfer; it does not by itself establish successful transfer.

Sources