Rosehip Mole asked

Why does random weight noise work on big models?

9 steps, drawn by Nib. Read it here, or watch him draw it.

1the puzzle

Normally we finetune a huge model with careful gradient steps. This paper asks a weirder question: what if you just jiggled the weights randomly instead?

2the cloud

Picture the pretrained weights as one point in a huge space. Around that point is a whole cloud of nearby parameter settings.

3small model

In a small model, task experts hide in a tiny sliver of that cloud — like needles in a haystack. You need gradient descent to find them.

4large model

But in a large, well-pretrained model, that neighborhood is packed — task-solving experts are everywhere nearby. Random noise practically trips over one.

5sample N

That density claim gives a wild recipe: sample N random perturbations of the weights, no gradients, no backprop at all.

6pick top K

Test each one on the task, keep only the top K performers — the accidental experts that random luck happened to land near.

7majority vote

Then don't pick a winner — run all K experts on new inputs and let them vote. Diversity plus majority vote beats any single noisy model.

8the payoff

The surprise: this dumb, fully parallel, gradient-free method is competitive with PPO, GRPO, and evolution strategies on today's large models.

9the thicket

So the real insight isn't the algorithm — it's the geometry. Pretraining at scale doesn't give you one good point, it gives you a thicket of good points.

Scale changes the shape of the search itself — with a big enough pretrained model, you barely need to search at all, just to look around.