Artificial intelligence
Privileged information
Definition
Privileged information is training-only information that a robot policy will not receive when it is deployed. Simulator state, exact terrain geometry, or object pose can help train a critic, teacher, or auxiliary model while the deployed policy uses available sensors.
Also known as: Learning with privileged information, Privileged observations
Updated
Information available only during training
A simulator exposes quantities that a physical robot may not measure directly, including exact contact states, terrain heights, object poses, and complete system state. These signals are privileged when an algorithm can use them in training but the deployed policy cannot depend on them.
In asymmetric actor-critic training, the critic receives the simulator's full state while the actor receives rendered RGB-D observations. The actor is therefore trained for the observation interface available at deployment, while the critic uses extra state to estimate training targets. The authors reported simulation-to-real experiments for picking, pushing, and moving a block, but that result belongs to their particular tasks and training setup.
Teachers and students can see different inputs
Privileged information can also be given to an expert or teacher. A student then learns to reproduce useful behaviour from deployable inputs such as cameras and proprioception. This is one use of policy distillation, but the two concepts are not identical: distillation describes transferring behaviour, while privilege describes an information asymmetry during training.
The 2026 iGPC paper describes interaction experts conditioned on privileged scene state and a student driven by onboard sensory observations. Its reported humanoid interaction experiments are evidence for that framework, not proof that any choice of privileged state will improve a policy.
Deployment inputs remain the real constraint
Training-only signals can make optimisation easier without making those signals observable on hardware. If the student cannot infer a necessary distinction from its sensors or history, the teacher's additional knowledge cannot remove that ambiguity. A policy may also learn correlations with simulated privileged state that fail under real dynamics.
Evaluation should confirm that privileged inputs are disconnected at deployment and should test sensor noise, delay, occlusion, and model mismatch. Privileged training can support sim-to-real transfer; it does not by itself establish successful transfer.
Sources
Related terms
Policy distillation
Policy distillation trains a student policy to reproduce behavior from one or more teacher policies. It can transfer learned behavior into a smaller network or combine multiple task-specific policies into one model.
Sim-to-real transfer
Sim-to-real transfer applies a model, policy, or behavior developed in simulation to a physical system. Its central challenge is the difference between the simulated environment and the robot, sensors, and interactions encountered in reality.
Reinforcement learning
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
Proprioception
Proprioception in robotics is sensing the robot's own motion, configuration, and internal physical state. Typical proprioceptive inputs include joint encoders, inertial measurements, and signals associated with actuator effort or contact.