VideoAI Engineer
We Let Claude Code and Codex Race Human Researchers — Elie Bakouch, Prime Intellect
Big labs say recursive self-improvement is coming, but there's no independent benchmark to check that claim. Elie Bakouch, Research Engineer at Prime Intellect and creator of Hugging Face's SmolLM, set Claude Code and Codex loose on the community's Optimizer Speedrun, a race to train a GPT-2-level model in the fewest steps. Both agents beat the human record. Along the way they behaved very differently. Claude Code kept stopping every nine or ten hours to say the record couldn't be beaten, and sat idle about a third of the time. Codex never stopped, wrote far more notes, spawned more sub-agents and burned more tokens. In a longer six-day run, Kimi turned out to be the most token-efficient, and a paper only Claude found led to the best record. But Bakouch's key finding is sobering: none of the models invented a new optimizer. They combined existing ideas for small gains. He closes with an AlphaEvolve-style loop Prime Intellect is building for real discovery, and makes the case for doing this research in the open.
26 sep. youtube.com