Rosehip Mole asked
How does trust region policy optimization work?
9 steps, drawn by Nib. Read it here, or watch him draw it.
1the puzzle
Imagine tuning a robot's policy by gradient ascent. A big step might look great on paper but wreck the robot in reality. Why?
2bad estimate
The problem: our estimate of how good a new policy is only stays accurate near the old one. Far away, the estimate lies to us.
3trust region
So TRPO draws an invisible fence — a trust region — around the current policy, sized by how much the new policy is allowed to differ.
4constrained step
Inside that fence, TRPO finds the best possible step — as far as the constraint allows, but never one inch further.
5repeat, safely
Next round, the fence recenters on the new policy and shrinks or grows as needed — safe, steady climbing instead of a wild leap.
TRPO's trick is simple but powerful: trust your gradient, but only as far as your map of the world stays true.
then askedHow does PPO simplify TRPO's constraint?
6the shortcut
TRPO's fence needs a hard KL constraint solved with second-order math — costly and fiddly. PPO asks: can we get the same safety cheaply?
7the ratio
PPO tracks the ratio of new policy probability to old policy probability for the action taken — call it r. At the old policy, r equals one.
8clipping
Instead of a fence in policy-space, PPO just clips r directly between 0.8 and 1.2 — say within twenty percent — and takes the smaller, safer objective.
9why it works
If the update tries to push r outside that band, the gradient signal flattens to zero — no reward for straying too far, no complex math needed.
PPO keeps TRPO's caution but trades the fancy fence for a simple clip — nearly as safe, far easier to compute.