Skip to content

Benchmarks

Which model is stronger: benchmark results, disputes over evaluation methods and leaderboard changes, on record.

9selectedRelated topicsModel releasesResearchReasoning

Latest selected

1–9 of 9

Oct 1

ThursdayToday

Sep 29

Tuesday

Sep 23

Wednesday

Aug 28

Friday

Aug 27

Thursday
  1. Piloting the world's first double-blind AI evaluations

    Google DeepMind announced the world’s first double-blind evaluation of proprietary frontier AI models, restricting external evaluations to cryptographically isolated environments to prevent models seeing test questions in advance. The pilot partners with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to test Gemini Flash Lite with confidential benchmarks in a privacy-preserving environment.

Aug 21

Friday
  1. Measuring benchmark optimization in speech recognition

    Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.

Aug 4

Tuesday

Jul 15

Wednesday

May 16

Saturday