The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI
Developers using autonomous agents wrote 741% more code but shipped only 30% more software.
The bottleneck is human review, and "just review harder" doesn't scale. Reviewer effectiveness collapses past about 400 lines, and agents now open 10,000-line PRs. Laurie Voss (Head of Developer Relations at Arize AI, co-founder of npm) reviews what the industry is actually doing about it. The evidence covers OpenAI's zero-human-code product, METR's finding that about half of SWE-bench-passing PRs wouldn't be merged, and Cognition's FrontierCode (88% on SWE-bench Pro vs. 29% on real mergeability). She explains why a mergeability benchmark would immediately become a training signal for frontier models. She also covers how Cursor and GitHub run review in production, why multi-pass review and default suspicion cut false positives, and what Carlini's agent-built C compiler and Bun's million-line Zig-to-Rust port (13,044 unsafe blocks) reveal about taking humans out of the loop. Automated reviewers can be fooled by prompt injection that humans catch, which leaves production as the last reviewer standing.
Texten är källans egen beskrivning av publiceringen. Innehållet tillhör AI Engineer.
Mer att läsa
Server-Side Code Execution Tools for AI Agents, Compared
OpenRouter för 9 tim sedan
v0.40.0
Ollama för 9 tim sedan
Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
TechCrunch AI för 13 tim sedan
Can ‘super intelligence’ and a non-binding safety pact solve AI’s image problem?
TechCrunch AI för 13 tim sedan