Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Deep dive on LLM Inference at Scale — Harshul Jain, Audible & Tanmay Sah, Independent AI Researcher

The video explains how to efficiently run large language models on GPUs—a practical challenge many teams face.

Example: a typical model can require 42 GB of memory just to store intermediate results when 80 users query simultaneously, which is why servers crash. The workshop identifies bottlenecks (memory growing with longer queries, slow initial response, low throughput) and demonstrates solutions on two fronts: making the model itself more efficient through techniques like "flash attention," and smarter job distribution on the server through "paged attention" and batch processing. Finally, it compares two popular server engines (vLLM and SGLang), showing they are equivalent for normal use but vLLM is three to four times faster when the model makes its own decisions and branches.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read