Prime Intellect executed more than 100 independent runs using 10 models to determine how autonomous agents navigate complex optimization tasks. The participants were tasked with iterating on a 124M GPT training recipe by adjusting optimizer hyperparameters without internet access, with the most successful run from Fable 5 narrowing the gap to a human record by 82%.

The experiments ran for up to 8 days on 8xH200 GPUs in a sandboxed environment, featuring models such as GPT-5.6 Sol, Kimi K3 and Grok 4.6. Prime Intellect released the full dataset, including reasoning streams and scratchpads, to allow the community to examine how models approach autonomous research. Elie Bakouch, a contributor to the project, noted that the setup scaled runtime and compute beyond similar tasks on OpenAI and Anthropic system cards, which often ran for less than a day.

Sign in to suggest edits

Key sources

  1. SOURCE@primeintellect“100+ autonomous runs across 10+ models, sandboxed on 8xH200s for up to 8 days”x.com
  2. SUPPORT@primeintellect“iterate on a 124M GPT training recipe from a shared baseline, only changing optimizer related hyperparameters, no internet access”x.com
  3. SUPPORT@primeintellect“Some even built small simulations to isolate a mechanism before deciding if another GPU run was worth it”x.com
  4. SOURCE@primeintellect“We release everything: full traces, scratchpads, reasoning streams from open-weight models, and our experiment setup”x.com
  5. SUPPORT@eliebakouch“this experiment is quite noisy, one run in the same setting has a ~50 step spread after 24h”x.com
Markdown