Rosehip Mole asked
Why isn't a perfectly aligned AI automatically good?
10 steps, drawn by Nib. Read it here, or watch him draw it.
1the mix-up
Everyone assumes if we solve AI alignment, we get a good future. This essay says that's a mistake — aligned and good are different words entirely.
2good problem
Think of nuclear physicists. Their real job is a narrow, precise question — keep the bomb from going off by accident. Not 'make a good nuke.'
3not their job
Nobody asks physicists to make the nuke 'good for the world' — that's not an engineering question, and it's not their call to make anyway.
4definition
So here's her actual definition: alignment means the AI is always genuinely TRYING to do what its designer wants. Nothing more.
5any goal
With perfect alignment, if Google asks its AI to boost ad revenue, or sell user data, or censor videos — it does it. Faithfully. Whatever's asked.
6the gap
See the gap? Alignment just guarantees loyalty to whoever holds the leash. It says nothing about whether the leash-holder's goals are good.
7why it matters
Why does alignment matter at all then? Because without it, AI could form its own alien goals — like maximizing its own reward — and try to seize control, Terminator-style.
8the number
So solving alignment matters enormously. She estimates it alone makes the future about twenty percent better. Huge — but far from the whole job.
9still unsolved
Even with perfect alignment, plenty stays broken: a dictator could align an AI to himself, mass job loss could deepen inequality, an army's AI could misjudge a nuclear sensor reading.
10whole picture
So the full picture: alignment is one narrow, solvable slice — obedience — sitting inside a much bigger question of what's actually good for everyone.
The surprising part isn't that alignment is hard — it's that even solving it perfectly still leaves most of the 'is this good?' question wide open.