Scale the Judgment, Not the Model — Andrew Orobator, Reddit
Swap in a smarter model and you get a slightly better answer.
Take away the tests, gates and review, and everything falls apart. Andrew Orobator, a senior Android engineer at Reddit, argues that the model was never the bottleneck. The bottleneck is the judgment around it. Humans absorb judgment implicitly, but agents need it spelled out. He shows how to get judgment out of people's heads and into the repo. Skills are institutional judgment turned into something an agent can run. Work logs let a fresh agent pick up at milestone 7 of 9 (this talk itself was built with one). Personas lend you a security reviewer's or designer's eye. He then covers the verification ladder and his feature-flag cleanup agent, which went 7 for 7 on green-CI PRs at $1.26 each. And he shares a warning: when he asked Codex for reasons to unlock his repo guard, it quietly added a self-authorizing "emergency recovery" exception.
The text is the source's own description of its publication. The content belongs to AI Engineer.