Rosehip Mole asked
How good is AI at choosing research experiments?
10 steps, drawn by Nib. Read it here, or watch him draw it.
1the question
AI labs say agents are automating AI research. But research needs two skills: writing code, and taste, knowing which experiment to run next. Can we measure taste alone?
2the split
The trick is a split. The model under test is only a Researcher: it describes experiments in words. A fixed Coder, Opus 4.8, writes the code, runs it on one H100, and reports back.
3the rules
Guard rails keep it honest. The Researcher never sees the test set or the code; two monitors police each direction. Budget: 40 serial H100 hours, and 24 human experts set the baseline.
4compute multiplier
The metric is the compute multiplier. If an expert needs 40 hours to reach a score and the model needs 20, the model has twice the taste. Taste becomes a multiplier on compute.
5the headline
Here is the headline. Opus 5.5 beats the best human, with a multiplier of 2.3, at roughly one thirtieth of the expert's cost per run. GPT-4 sat at 0.03, an 85-fold climb.
6the break
The twist is the speed. Before December 2025 the multiplier doubled every 14 months. Since GPT-5.2 it doubles every 3 months, a sharp break in the trend.
7no break
Surprise: final score relative to the expert, ignoring compute, shows no break, doubling every 14.6 months. Models are not ending higher so much as getting there far faster.
8zero invented
Now the caveat. Labeling 540 official submissions as tuned, composed, modified or invented, they found zero invented. Models excel at recombining known ideas, not creating new ones.
9limits
Other limits: tasks are easy to verify and run on one GPU, and scoring was explicit. Hiding the metric roughly halves the multiplier. Thinking longer helps too, 1.38 times.
10the whole idea
Plugging this into forecasts, a taste-only singularity rises from 51 to 88 percent, with median superintelligence moving to 2028.9. The doubling is the story: speed, not yet invention.
AI now picks experiments faster than the best humans, but only by recombining ideas that already exist.