Tokens Should Have Jobs — Katelyn Lesse & Angela Jiang, Anthropic
Researchers from Anthropic show that giving AI agents larger token budgets (tokens are the building blocks of AI responses) isn't always the answer.
With a fixed budget of 600,000 tokens, an agent that only executed tasks performed worse than one that used tokens more strategically: by consulting an adviser, by having a grader review results iteratively, or by having the agent study its own previous attempts. For financial analysis, an answer that is 80 percent correct often isn't useful — one wrong figure ruins the whole result. A simple execution succeeded in only 42 percent of tests on the first try, costing roughly 1.8 million tokens total. Smarter strategies could reach the same quality with fewer tokens.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.