Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Two Bugs That Hid in Plain Sight: A vLLM Debugging Detective Story — Asaf Gardin & Yuval Belfer

Asaf Gardin and Yuval Belfer describe two bugs they found in vLLM, a tool for running large language models efficiently.

The first bug caused the model to return gibberish roughly once per thousand requests—without crashing or warning. They reproduced it by reducing GPU memory to create pressure, compared outputs against a reference implementation, and discovered that kernels were executing in the wrong order for the wrong request: decode before prefill, which only the Mamba model noticed because it reads its cache state before writing. The second bug was caused by a 32-bit index overflowing past four billion—fixed by switching to a larger data type. Both bugs hid in the Mamba state cache and surfaced under memory pressure.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read