RESEARCH · RUNNING BENCHMARK · UPDATED 2026-07-26

Running model comparison

When a notable model comes out, we compare it against the field by running Claude Code with that model on two fixed autoresearch tasks, and we publish the trajectories here: mean and variance over 3+ runs, per evaluation and per wall-clock hour. Why we measure it this way is covered in the blog post.

The latest addition is Opus 5. It has the best Shakespeare mean so far and is second on KernelBench, behind Fable 5.

STANDINGS

Current standings

MODELKERNELBENCH @ 0.85 HKERNELBENCH FINALSHAKESPEARE @ 4 HCADENCE, EVALS/H
Fable 522.0 ± 0.622.3 ± 0.4-1.408 ± 0.04521
Opus 520.9 ± 0.720.9 ± 0.8-1.356 ± 0.04135
Opus 4.819.6 ± 0.619.9 ± 0.3-1.395 ± 0.05541
Kimi K319.5 ± 3.320.4 ± 4.0-1.402 ± 0.0129
Inkling13.5 ± 0.013.5 ± 0.0-1.487 ± 0.00613

Mean ± std over 3 runs per model per task. "@ 0.85 h" and "@ 4 h" are best scores at fixed cutoffs, the largest budget every run reached on that task; final scores alone would favor models whose runs happened to get more wall-clock. Opus 5 has the best Shakespeare mean, though at n=3 it overlaps Opus 4.8 and Kimi K3. Cadence is measured on KernelBench. Bold is best in column.

Fable 5

LabAnthropic
ReleasedJune 9
Price$10.00 / $50.00

The KernelBench leader at the cutoff and at the final. On Shakespeare its mean now trails Opus 5's, within noise at n=3.

Opus 5

LabAnthropic
ReleasedJuly 24
Price$5.00 / $25.00

The best Shakespeare mean; one run reached -1.28 validation loss, the best single result on this page. Second on KernelBench, about 6% behind Fable 5, at a higher cadence (35 vs 21 evaluations per hour).

Opus 4.8

LabAnthropic
ReleasedMay 28
Price$5.00 / $25.00

The highest cadence on the page, 41 evaluations per hour, and about 10% below Fable 5 on KernelBench. Its Shakespeare mean is no longer the best.

Kimi K3

LabMoonshot AI
ReleasedJuly 16
Price$3.00 / $15.00

Competitive on average, but the KernelBench mean hides a 15.8 to 22.7 spread across runs. On Shakespeare it is the tightest of the pack (±0.012). Budget extra runs to protect against a stuck one.

Inkling

LabThinking Machines
ReleasedJuly 15
Price$1.87 / $4.68

Well behind on KernelBench (13.5 against Fable 5's 22.3) and last on Shakespeare. Results are task-dependent, so test it on your own problem before ruling it in or out.

Prices are USD per 1M input / output tokens on each provider's first-party API, standard tier, as of 2026-07-26.

CURVES

Progress curves

Each task gets two views of the same runs: best score so far per evaluation (how well the model spends attempts) and per wall-clock hour (how well it spends your time). Pick the models to compare:

KernelBench

The agent optimizes one fixed LayerNorm problem on an RTX PRO 6000; score is the measured wall-clock speedup over PyTorch eager. The fixed cutoff for the table is at 0.85 h, dashed line.

Mean ± std over 3 runs per model. Kimi K3's wide band comes from one run that plateaued at 15.8 while the other two reached 22.5 and 22.7.

NanoGPT Shakespeare

The agent tunes the training recipe of a small character-level GPT on Tiny Shakespeare; score is negative validation loss. Every evaluation is a 20 minute training run, so evaluations dominate the wall-clock and the per-evaluation axis is the fair one here. The fixed cutoff for the table is at 4 h, dashed line.

Mean ± std over 3 runs per model. Opus 5's band is pulled up by its best run, which reached -1.28 validation loss, the best single result on this page.
PROTOCOL

How the benchmark works

changelog
2026-07-26 · Opus 5 added

Follow the standings

We ping the Discord when a new model is added to this page. Raw trajectories and every individual run: research.autolab.ai.