Artificial intelligence
Hierarchical reinforcement learning
Definition
Hierarchical reinforcement learning organizes learned decision-making into levels, often with a higher-level policy selecting goals or skills and lower-level policies producing actions. The levels can operate over different time scales.
Also known as: HRL
Updated
Separate goals from detailed actions
The HIRO research studies a hierarchy in which a higher-level controller proposes learned goals and a lower-level controller learns behavior conditioned on those goals.
For a mobile manipulator, a conceptual hierarchy might first choose a nearby location to reach and then produce the movements needed to reach it. This is an example of the division of responsibility, not a claim that HIRO demonstrated that exact physical platform.
Levels are learned together
A hierarchy can reduce the burden on one policy to decide both long-term progress and every immediate action. HIRO uses off-policy experience to train both levels and evaluates complex behaviors in simulated robotic tasks.
The method also illustrates a difficulty: when the lower-level policy changes, the same high-level goal can lead to different behavior. HIRO introduces a correction to account for this change when reusing past experience.
Hierarchy alone is not reinforcement learning
A hand-written task tree or a language model that calls fixed robot skills is hierarchical, but that alone does not make it hierarchical reinforcement learning. The term concerns reward-based learning within the hierarchy.
The chosen goals, skill interfaces, and termination conditions determine what the levels can express. A useful hierarchy for one robot task may not provide the right structure for a different task or body.
Sources
Related terms
Reinforcement learning
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
Language-conditioned policy
A language-conditioned policy selects actions using a language instruction together with observations. The instruction specifies or modifies the behavior requested from the policy.
Generalist robot policy
A generalist robot policy is a learned action-selection model designed to perform multiple tasks across a range of robot settings. Its generality depends on the tasks, observations, action interfaces, and robot bodies included in training and evaluation.