Context Engineering in 2026 — Louis-François Bouchard, Omar Solano & Samridhi Vaid, Towards AI
In a video from Towards AI, Louis-François Bouchard, Omar Solano, and Samridhi Vaid present results from experiments on how to best handle long context history in AI systems.
They tested an AI tutor and found that it is often better to keep all history instead of trying to compress it — because prompt caching (where APIs save tokens and reuse them more cheaply) makes compression actually more expensive when it destroys the cache. In their measurements, the uncompressed version retained details 95 percent of the time compared to 32 percent after summarization, and was able to maintain specific facts up to 800,000 tokens. The main rule they conclude is to first identify which constraint you actually have — context window size, speed, or cost — before you start compressing.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.