Skip to content
VibekollenBETAVibekollen
VideoAI Engineer

AI Evals for Cross-Functional Teams — Nachiket Paranjape & Swaroop Chitlur Haridas, DoorDash

DoorDash builds its AI evaluation (evals) by allowing non-technical teams to create their own tools for testing and improving AI.

Rather than building a perfect interface from the start, the platform team created stable APIs — standardized connection points — that strategy, operations, and product teams could plug their own tools into. Evaluation became a multi-team workflow: operations staff annotate and label data, product managers set quality requirements, and engineers provide measurement data. The process itself is straightforward — they collect AI responses, mark the best examples, train a "judge" (a small AI model that grades automatically), and monitor performance over time. Since different teams have different needs, they accept that there is no single right way, and the cost per annotation job dropped significantly.

Open on YouTube →

Vibekollen prepared this summary with AI from the original publication. The content belongs to AI Engineer.

More to read