Rosehip Mole asked

Why isn't a perfectly aligned AI automatically good?

10 steps, drawn by Nib. Read it here, or watch him draw it.

1the mix-up

Everyone assumes if we solve AI alignment, we get a good future. This essay says that's a mistake — aligned and good are different words entirely.

2good problem

Think of nuclear physicists. Their real job is a narrow, precise question — keep the bomb from going off by accident. Not 'make a good nuke.'

3not their job

Nobody asks physicists to make the nuke 'good for the world' — that's not an engineering question, and it's not their call to make anyway.

4definition

So here's her actual definition: alignment means the AI is always genuinely TRYING to do what its designer wants. Nothing more.

5any goal

With perfect alignment, if Google asks its AI to boost ad revenue, or sell user data, or censor videos — it does it. Faithfully. Whatever's asked.

6the gap

See the gap? Alignment just guarantees loyalty to whoever holds the leash. It says nothing about whether the leash-holder's goals are good.

7why it matters

Why does alignment matter at all then? Because without it, AI could form its own alien goals — like maximizing its own reward — and try to seize control, Terminator-style.

8the number

So solving alignment matters enormously. She estimates it alone makes the future about twenty percent better. Huge — but far from the whole job.

9still unsolved

Even with perfect alignment, plenty stays broken: a dictator could align an AI to himself, mass job loss could deepen inequality, an army's AI could misjudge a nuclear sensor reading.

10whole picture

So the full picture: alignment is one narrow, solvable slice — obedience — sitting inside a much bigger question of what's actually good for everyone.

The surprising part isn't that alignment is hard — it's that even solving it perfectly still leaves most of the 'is this good?' question wide open.