Skip to content
Google DeepMind·· 2026-08-27

Piloting the world's first double-blind AI evaluations

Piloting the world's first double-blind AI evaluations

AI summary

Google DeepMind announced the world’s first double-blind evaluation of proprietary frontier AI models, restricting external evaluations to cryptographically isolated environments to prevent models seeing test questions in advance. The pilot partners with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to test Gemini Flash Lite with confidential benchmarks in a privacy-preserving environment.

Selection record

AdmittedSum of both 120 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:双盲AI模型评测,实质AI评测技术
Why it was chosen
The double-blind Gemini Flash Lite evaluation in a cryptographically isolated environment demonstrates a technical approach to benchmark contamination.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Google DeepMind · deepmind.google