Artificial intelligence

Hierarchical reinforcement learning

Definition

Hierarchical reinforcement learning organizes learned decision-making into levels, often with a higher-level policy selecting goals or skills and lower-level policies producing actions. The levels can operate over different time scales.

Also known as: HRL

Updated

Separate goals from detailed actions

The HIRO research studies a hierarchy in which a higher-level controller proposes learned goals and a lower-level controller learns behavior conditioned on those goals.

For a mobile manipulator, a conceptual hierarchy might first choose a nearby location to reach and then produce the movements needed to reach it. This is an example of the division of responsibility, not a claim that HIRO demonstrated that exact physical platform.

Levels are learned together

A hierarchy can reduce the burden on one policy to decide both long-term progress and every immediate action. HIRO uses off-policy experience to train both levels and evaluates complex behaviors in simulated robotic tasks.

The method also illustrates a difficulty: when the lower-level policy changes, the same high-level goal can lead to different behavior. HIRO introduces a correction to account for this change when reusing past experience.

Hierarchy alone is not reinforcement learning

A hand-written task tree or a language model that calls fixed robot skills is hierarchical, but that alone does not make it hierarchical reinforcement learning. The term concerns reward-based learning within the hierarchy.

The chosen goals, skill interfaces, and termination conditions determine what the levels can express. A useful hierarchy for one robot task may not provide the right structure for a different task or body.

Sources