Routing LLM Inference in Production: From Engine Signals to Policy — Qianru Lao & Lu Zhang, OpenAI
OpenAI has developed a new routing system for its AI inference in production.
Previously they used a feedback loop where engines reported performance signals that a controller converted to weights, but this created hard-to-explain behavior and oscillation that damaged cache performance. The new system uses a control plane with global visibility of all clusters and a data plane in each cluster that selects which engine serves each request based on cached weights. An optimizer minimizes end-to-end latency under the constraint that all requests are routed and no engine is overloaded, with examples of how geographic proximity is not always optimal — a slow response from the nearest engine can lose to faraway engines.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.