Skip to content
VibekollenBETAVibekollen
BlogGoogle DeepMind

Piloting the world's first double-blind AI evaluations

Article image or reusable cover for Google DeepMind

Google DeepMind introduces the world's first "double-blind" evaluation of AI models — a method where evaluators and model creators use cryptographic technology to test models without either party seeing the other's sensitive data.

The problem being solved is that AI models can be trained on the same test questions used to assess them, making results unreliable. By using Google's Confidential Computing, evaluators can test the Gemini Flash Lite model without Google seeing the test questions and without the model being able to optimize itself based on them in advance.

If a model has already seen the test questions - a problem known as benchmark contamination - the results can only be trusted to an extent.
Verbatim from the article at Google DeepMind
Read the full story at Google DeepMind →

Vibekollen prepared this summary with AI from the original publication. The content belongs to Google DeepMind.

More to read