Skip to content
Hugging Face Blog·· 2026-08-21

Measuring benchmark optimization in speech recognition

Measuring benchmark optimization in speech recognition

AI summary

Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.

Selection record

AdmittedSum of both 120 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:ASR基准优化研究,涉AI模型评测
Why it was chosen
Three probes quantify speech models' benchmark overfitting, with reusable detection scripts and access to the leaderboard.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Hugging Face Blog · huggingface.co