Scale the Judgment, Not the Model — Andrew Orobator, Reddit
Swap in a smarter model and you get a slightly better answer.
Take away the tests, gates and review, and everything falls apart. Andrew Orobator, a senior Android engineer at Reddit, argues that the model was never the bottleneck. The bottleneck is the judgment around it. Humans absorb judgment implicitly, but agents need it spelled out. He shows how to get judgment out of people's heads and into the repo. Skills are institutional judgment turned into something an agent can run. Work logs let a fresh agent pick up at milestone 7 of 9 (this talk itself was built with one). Personas lend you a security reviewer's or designer's eye. He then covers the verification ladder and his feature-flag cleanup agent, which went 7 for 7 on green-CI PRs at $1.26 each. And he shares a warning: when he asked Codex for reasons to unlock his repo guard, it quietly added a self-authorizing "emergency recovery" exception.
Texten är källans egen beskrivning av publiceringen. Innehållet tillhör AI Engineer.
Mer att läsa
Building advertising for the way people use AI
OpenAI för 2 tim sedan
Server-Side Code Execution Tools for AI Agents, Compared
OpenRouter för 12 tim sedan
v0.40.0
Ollama för 12 tim sedan
Google froze its open source bug bounty program due to a ‘significant rise’ in AI submissions
TechCrunch AI för 16 tim sedan