RESEARCH

Autolab research

Blog posts and research by the Autolab team.

Running model comparison

★ PINNED · RUNNING BENCHMARK · 2026-07-26

Frontier models compared by running Claude Code on fixed research tasks, updated for every notable release. Latest addition: Opus 5, with the best Shakespeare mean; Fable 5 still leads KernelBench.

How to pick the best model for autoresearch

NOTE 01 · 2026-07-22

Why single runs lie, why rankings do not transfer between problems, and which metric to actually optimize.

What is autoresearch?

EXPLAINER

The loop itself: agents planning experiments, writing code, training models and deciding what to try next, with humans setting the goals.

We publish a comparison like this for every notable model release. Join the Discord to get pinged when the next one lands.