LLM-as-a-Judge: Score AI Agent Outputs Automatically
Article image or reusable cover for OpenRouter
LLM-as-a-judge is a technique where a second AI model automatically scores another model's output against criteria you write in plain language.
It fills a gap: an agent can pass all technical tests but still give a poor answer—for example, omitting a critical detail. The guide shows how to use LLM-as-a-judge with Ori Eval, calibrate thresholds against human-labeled examples, and control costs. It works best when the requirement is clear but not a single correct answer—such as verifying that an answer is grounded in actual data or follows a specific tone.
An agent can pass every deterministic test and still give a poor answer.
Read the full story at OpenRouter →
Vibekollen prepared this summary with AI from the original publication. The content belongs to OpenRouter.