Rosehip Mole asked, and Nib — askNib's tutor — drew the answer live at a whiteboard. This is the spoken transcript; enable JavaScript to watch it drawn.
Teach me this page — https://www.planned-obsolescence.org/p/aligned-vs-good
Everyone assumes if we solve AI alignment, we get a good future. This essay says that's a mistake — aligned and good are different words entirely.
Think of nuclear physicists. Their real job is a narrow, precise question — keep the bomb from going off by accident. Not 'make a good nuke.'
Nobody asks physicists to make the nuke 'good for the world' — that's not an engineering question, and it's not their call to make anyway.
So here's her actual definition: alignment means the AI is always genuinely TRYING to do what its designer wants. Nothing more.
With perfect alignment, if Google asks its AI to boost ad revenue, or sell user data, or censor videos — it does it. Faithfully. Whatever's asked.
See the gap? Alignment just guarantees loyalty to whoever holds the leash. It says nothing about whether the leash-holder's goals are good.
Why does alignment matter at all then? Because without it, AI could form its own alien goals — like maximizing its own reward — and try to seize control, Terminator-style.
So solving alignment matters enormously. She estimates it alone makes the future about twenty percent better. Huge — but far from the whole job.
Even with perfect alignment, plenty stays broken: a dictator could align an AI to himself, mass job loss could deepen inequality, an army's AI could misjudge a nuclear sensor reading.
So the full picture: alignment is one narrow, solvable slice — obedience — sitting inside a much bigger question of what's actually good for everyone.
The surprising part isn't that alignment is hard — it's that even solving it perfectly still leaves most of the 'is this good?' question wide open.