Vertical Mobility: Inference from MVP to Trillion-Parameter Workloads — Sitanshu Gupta, CoreWeave
Sitanshu Gupta from CoreWeave describes how to build an inference stack—the system that runs AI models—that efficiently handles vastly different types of work.
The key insight is that when users send multiple questions in sequence, much of the information is identical, making it far more expensive to process the first time than later. CoreWeave offers three payment models: serverless where you pay per token without seeing hardware, a middle option for customers who know their traffic patterns, and dedicated where you control everything. The system matches four different workload patterns—chat questions, agent decisions, voice and video, and batch jobs—together like Tetris to maximize hardware utilization. Two major optimizations are compressing model weights to four bits and training fast predictors on a customer's own data to reduce latency.
Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.