KV Cache-Aware Routing and P/D Disaggregation on Kubernetes — Yuchen Fama & Ashish Kamra, Red Hat
Red Hat engineers Yuchen Fama and Ashish Kamra present two methods to accelerate AI model execution on servers.
The first is intelligent routing — when users ask follow-up questions, previous computations are reused on the same server, making responses 3 times faster. The second is dividing work between two server types: one that prepares text and one that generates words. In a practical test, latency between words dropped from 900 milliseconds to 100 milliseconds. However, disaggregation only works well under moderate load and requires specialized network hardware — otherwise it is better to keep everything on a single server.
Open on YouTube →
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.