Hoppa till innehåll
VibekollenBETAVibekollen
VideoAI Engineer

The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI

Developers using autonomous agents wrote 741% more code but shipped only 30% more software.

The bottleneck is human review, and "just review harder" doesn't scale. Reviewer effectiveness collapses past about 400 lines, and agents now open 10,000-line PRs. Laurie Voss (Head of Developer Relations at Arize AI, co-founder of npm) reviews what the industry is actually doing about it. The evidence covers OpenAI's zero-human-code product, METR's finding that about half of SWE-bench-passing PRs wouldn't be merged, and Cognition's FrontierCode (88% on SWE-bench Pro vs. 29% on real mergeability). She explains why a mergeability benchmark would immediately become a training signal for frontier models. She also covers how Cursor and GitHub run review in production, why multi-pass review and default suspicion cut false positives, and what Carlini's agent-built C compiler and Bun's million-line Zig-to-Rust port (13,044 unsafe blocks) reveal about taking humans out of the loop. Automated reviewers can be fooled by prompt injection that humans catch, which leaves production as the last reviewer standing.

Öppna på YouTube →

Texten är källans egen beskrivning av publiceringen. Innehållet tillhör AI Engineer.

Mer att läsa