Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

From Ingestion to Agents: How AI Teams Build on Document Intelligence — Adit Abraham, Reducto

The newest frontier model scores about thirty percent on a data lab's benchmark of decisions from PDFs, and Adit Abraham has met people who worked on PDF processing before he was born.

The format was built to print, not to be reasoned over, and humans encode meaning visually: merged cells, line charts, unreadable handwriting. Reducto has processed billions of them, and the talk is the lessons, not the product. RAG meant a bad parse cost one answer; with agents, bad inputs compound across every step. VLMs finally read the long tail like a human, but they are not one size fits all: small detectors still find layout on a CPU at scale, and a VLM asked to rewrite OCR will helpfully recompute a total the human got wrong. His agentic OCR applies token level corrections, a zero for an O, instead of regenerating the page. Simple tables go to markdown and complex ones to HTML, but embedding models cannot match how did revenue change to a blob of tags, so a natural language rendering serves retrieval. Parsed structure rather than raw PDFs lifted other frontier models past the newest one on that benchmark and cut reasoning tokens.

Open on YouTube →

The text is the source's own description of its publication. The content belongs to AI Engineer.

More to read