Google Research·· 2026-04-01
Building better AI benchmarks: How many raters are enough?
Building better AI benchmarks: How many raters are enough?
AI summary
Google Research studied trade-offs between evaluation item counts and annotators per item in Forest vs Tree: The (N, K) Trade-off in Reproducible ML Evaluation. Common counts of one, three or five annotators are often insufficient; reflecting human disagreement frequently requires more than ten.
Selection record
Threshold 60Official, first-handFirst 31Second 31
Not admittedSum of both 62 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:研究AI基准测试中评分者数量与可复现性
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Google Research · research.google