Supercharge your

A thousand experiments in parallel, on your compute, scored against your evals. State of the art with the team and budget you already have.

$ curl -fsSL app.autolab.ai/install.sh | sh
autolab · from goal to merge
a real session, compressed

One thousand agents. One goal.

Write down what you want: beat this baseline, cut this latency, test these three ideas. Autolab reads your repo, plans the experiments, and spins up an agent for each one, across every machine you give it. They keep going until the goal is met.

autolab · your-repo/frontier-model24 AGENTS · LIVE
agent-01lr-sweep
sweep 1e-4 → 3e-4
+0.4%
agent-02data-mix
hard-mix v2 · 60/40
+0.8%
agent-03rotary-θ
θ = 1e6 · long ctx
+1.1%
agent-04tokenizer
vocab 64k sweep
−0.1%
agent-05muon-opt
✕ NaN loss · gradients exploded
running
agent-06quant-int8
✕ OOM on node-04 · resharding
running
agent-07gqa-heads
8 → 4 kv heads
+0.2%
agent-08distill-7b
spawned by agent-03
queued
agent-09evals
reading 412 log lines
click an agent to inspect→ merged 9f41c2 · +2.3%

Understand what happened, then aim the next thousand.

Every experiment is a commit: code, environment, metrics, logs, checkpoints. The readout ranks what worked and why, so you always know what to try next.

BEST VAL_ACC OVER TIMElast 24h · 56 experiments91%92%93%94%24h agonowbaseline 91.2%data-mix v2+0.8%rotary θ = 1e6+1.1%full scale · 8×H100+0.4%93.5% · merged 9f41c2
full scale · 8×H100 · live
every cell is a GPU · nodes spin down and refill · always on

No GPU sits idle.

Idle compute is the most expensive thing you own. Autolab keeps your cluster saturated with the next-most-valuable experiment, and scales the winner up when it’s proven.

You name the metric.

Autolab optimizes whatever your eval measures. The agents push the number, and nothing ships unless it moves.

Higher accuracy

Architectures, data mixes, optimizers, schedules. Every promising variant gets its shot.

▲ +2.3% val_accSCORED AGAINST YOUR EVAL

Lower latency

Quantization, kernels, serving configs, tuned on the stack you actually serve from.

▼ −31% p99SCORED AGAINST YOUR EVAL

Lower cost

The cheapest training and serving configuration that still clears your quality bar.

▼ −38% cost / runSCORED AGAINST YOUR EVAL
numbers shown are illustrative targets · your eval defines the win

Start this afternoon.

Install, point it at your eval, go. Or meet Autolab inside the coding agent you already use.

CLI Claude Code Codex
$ curl -fsSL app.autolab.ai/install.sh | sh
$ autolab init   # point it at the eval script that prints your metric
$ autolab start  # the agents begin. watch or walk away.
$ autolab install claude-code
then, inside Claude Code:
> use /autolab to run experiments on this repo
$ autolab install codex
then, inside Codex:
> queue autolab experiments against eval.py

Your infra or ours.

On your cluster or your cloud account: code, data, and weights never leave your network. On-prem installs available for pilots.

Built by researchers who got tired of waiting.

RESEARCHERS FROM
Reality Labs
Careers →

Questions, answered.

What is autoresearch?+

Autoresearch is autonomous AI agents running the machine learning research loop: proposing experiments, writing the code, training and evaluating models, and deciding what to try next — while humans set the goal. Instead of one researcher hand-running one experiment at a time, an autoresearch platform runs thousands in parallel and merges only what improves the metric. Read the full explanation →

What is Autolab?+

Autolab is an autoresearch platform for AI model training. You give it a goal and an eval; its agents read your repo, plan experiments, launch them across your GPUs, score every run against your metric, and keep iterating until the goal is met. It was built by ML researchers from MIT, Harvard, Stanford, and Google DeepMind.

What is agentic training of models?+

Agentic model training means AI agents, not humans, drive the training loop: they sweep learning rates, change data mixes, modify architectures, launch runs, read the training logs, and decide the next experiment. Autolab coordinates hundreds of these agents against a single goal, on your own compute.

How is this different from AutoML or hyperparameter tuning?+

AutoML and hyperparameter optimization search a fixed space with predefined strategies. Autoresearch agents write real code in your repository: new data pipelines, loss functions, kernels, quantization and serving configs — anything a researcher could try — and every change is scored against your own evals before it merges.

Where do the experiments run?+

On your cluster or in your cloud account. Code, data, and weights never leave your network, and on-prem installs are available for pilots. Autolab keeps whatever GPUs you have — from a spare 3090 to a multi-node H100 cluster — saturated with the next-most-valuable experiment.

What can Autolab optimize?+

Whatever your eval measures: model accuracy, training cost, inference latency, throughput. Teams use it for pre-training and fine-tuning experiments, post-training recipes, and inference optimization — the agents push the number, and nothing ships unless it moves.

One researcher. A lab’s worth of insights.
We are entering the era of autonomous research.
Book a demo