RESEARCH
Autolab research
Blog posts and research by the Autolab team.
Running model comparison
★ PINNED · RUNNING BENCHMARK · 2026-07-26
Frontier models compared by running Claude Code on fixed research tasks, updated for every notable release. Latest addition: Opus 5, with the best Shakespeare mean; Fable 5 still leads KernelBench.
How to pick the best model for autoresearch
NOTE 01 · 2026-07-22
Why single runs lie, why rankings do not transfer between problems, and which metric to actually optimize.
What is autoresearch?
EXPLAINER
The loop itself: agents planning experiments, writing code, training models and deciding what to try next, with humans setting the goals.
We publish a comparison like this for every notable model release. Join the Discord to get pinged when the next one lands.