Artificial intelligence

Reward shaping

Definition

Reward shaping adds supplementary rewards to guide reinforcement learning toward useful behavior. Poorly chosen shaping can change which policy is optimal, so an easier training signal is not automatically equivalent to the original task objective.

Updated

Provide feedback before task completion

A robot may receive its main reward only after reaching a goal. Shaping can provide intermediate feedback, such as a signal related to progress toward that goal. Ng, Harada, and Russell analyze when adding such feedback preserves the original optimal policy.

The distinction matters because a learner optimizes the reward it actually receives, including any extra terms.

Potential-based shaping has a precise form

For a discounted problem, potential-based shaping adds gamma * Phi(next_state) - Phi(state), where Phi assigns a potential to each state and gamma is the discount factor. The paper establishes policy invariance under its stated Markov decision process assumptions and boundary conditions.

This form rewards a change in potential rather than handing out an unrelated bonus whenever a convenient event occurs.

Extra rewards can create unintended loops

The paper describes examples in which rewarding progress or repeated ball contact encouraged behavior that failed the intended objective. For a robot moving a part, repeatedly collecting a proximity bonus could likewise become preferable to completing placement if the reward is poorly specified.

Shaping is a tool within reinforcement learning, not a substitute for defining the task. Report the full reward and verify actual task completion rather than relying only on the shaped return.

Sources