Artificial intelligence
Reinforcement learning
Definition
Reinforcement learning trains an agent to choose actions that maximize expected cumulative reward through experience with an environment. In robotics, the learned policy can select movements or higher-level behaviors from observations.
Also known as: RL
Updated
Optimize outcomes over time
OpenAI's reinforcement-learning documentation describes an agent that takes actions, receives observations and rewards, and seeks a high return, or accumulated reward. A policy is the rule that selects actions; a value function estimates expected future return.
A robot's observation can include joint angles, velocities, or camera images. Its action space might contain continuous commands rather than a small list of discrete choices.
Reward and demonstration are different signals
A reward indicates how an outcome contributes to the task objective. It does not necessarily specify the exact action the robot should take. Imitation learning instead uses demonstrated behavior, although practical systems can combine both signals.
For example, an object-pushing task can reward getting an object to a target. The learner must discover an action sequence that achieves that outcome.
Training conditions shape the result
Physical interaction is not the only source of experience. Simulation-based robot-control research trains policies in simulated environments before evaluating transfer to a real arm.
Reward design, observations, available actions, and training dynamics all affect the learned behavior. High return in a simulator does not by itself establish reliable hardware performance, and a high reward does not establish that an incomplete task specification captured every desired constraint.
Sources
Related terms
Imitation learning
Imitation learning learns behavior from examples supplied by a demonstrator. In robotics, demonstrations can teach a policy how to perform a task without requiring every action or objective to be programmed by hand.
Offline reinforcement learning
Offline reinforcement learning learns a reward-optimizing policy from previously collected experience without gathering new environment interactions during that learning stage. The data may come from earlier policies, demonstrations, or other collection procedures.
Reward shaping
Reward shaping adds supplementary rewards to guide reinforcement learning toward useful behavior. Poorly chosen shaping can change which policy is optimal, so an easier training signal is not automatically equivalent to the original task objective.