Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Building the Document Context Layer for AI Agents — Jerry Liu, LlamaIndex

A PDF is designed for printing, so its text is glyphs with coordinates and its tables are line segments with characters at cell positions.

Reading order for a multicolumn page is not stored anywhere. That, Jerry Liu argues, is why an agent handed a raw file cannot make sense of it, and why document OCR is still unsolved after twenty years. His larger frame is that RAG in 2026 decomposes into an agent harness plus a context layer. Retrieval complexity has moved into the agent, which now reasons about the right search term instead of hacking around top k retrieval; context has moved up the stack toward MCP servers and skills; and programs are increasingly written in English. For LlamaIndex, what remains is the ten trillion plus pages locked in PDFs, PowerPoints, Word documents, and spreadsheets. The platform he describes has three layers: parsing into token efficient markdown and metadata, semantic storage as document management for humans and agents, and repeatable workflows for invoices, KYC, and claims.

Open on YouTube →

The text is the source's own description of its publication. The content belongs to AI Engineer.

More to read