Scaling to Long Horizons — Ross Taylor & Chengxi Taylor, General Reasoning
Ross Taylor and Chengxi Taylor from General Reasoning discuss training AI agents to maintain focus and perform well over extended periods—hours rather than seconds.
They explain techniques such as value models, which help AI understand which choices lead to long-term success, and bootstrapping, which extracts useful signals from limited rewards. They demonstrate how frontier models failed when given real money to trade football matches, revealing that the environment where the AI was trained was not realistic enough. The key insight is that succeeding with long horizons requires not just larger context windows but, most importantly, better and more simulated environments for training.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.