Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI

OpenAI has developed a new routing system for its AI inference in production.

Previously they used a feedback loop where engines reported performance signals that a controller converted to weights, but this created hard-to-explain behavior and oscillation that damaged cache performance. The new system uses a control plane with global visibility of all clusters and a data plane in each cluster that selects which engine serves each request based on cached weights. An optimizer minimizes end-to-end latency under the constraint that all requests are routed and no engine is overloaded, with examples of how geographic proximity is not always optimal — a slow response from the nearest engine can lose to faraway engines.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read