Rosehip Mole asked

How does trust region policy optimization work?

9 steps, drawn by Nib. Read it here, or watch him draw it.

1the puzzle

Imagine tuning a robot's policy by gradient ascent. A big step might look great on paper but wreck the robot in reality. Why?

2bad estimate

The problem: our estimate of how good a new policy is only stays accurate near the old one. Far away, the estimate lies to us.

3trust region

So TRPO draws an invisible fence — a trust region — around the current policy, sized by how much the new policy is allowed to differ.

4constrained step

Inside that fence, TRPO finds the best possible step — as far as the constraint allows, but never one inch further.

5repeat, safely

Next round, the fence recenters on the new policy and shrinks or grows as needed — safe, steady climbing instead of a wild leap.

TRPO's trick is simple but powerful: trust your gradient, but only as far as your map of the world stays true.

then askedHow does PPO simplify TRPO's constraint?

6the shortcut

TRPO's fence needs a hard KL constraint solved with second-order math — costly and fiddly. PPO asks: can we get the same safety cheaply?

7the ratio

PPO tracks the ratio of new policy probability to old policy probability for the action taken — call it r. At the old policy, r equals one.

8clipping

Instead of a fence in policy-space, PPO just clips r directly between 0.8 and 1.2 — say within twenty percent — and takes the smaller, safer objective.

9why it works

If the update tries to push r outside that band, the gradient signal flattens to zero — no reward for straying too far, no complex math needed.

PPO keeps TRPO's caution but trades the fancy fence for a simple clip — nearly as safe, far easier to compute.