Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat

Red Hat engineers Yuchen Fama and Ashish Kamra present two methods to accelerate AI model execution on servers.

The first is intelligent routing — when users ask follow-up questions, previous computations are reused on the same server, making responses 3 times faster. The second is dividing work between two server types: one that prepares text and one that generates words. In a practical test, latency between words dropped from 900 milliseconds to 100 milliseconds. However, disaggregation only works well under moderate load and requires specialized network hardware — otherwise it is better to keep everything on a single server.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read