Skip to content

#Benchmarks

1 today

Oct 1

ThursdayToday1 items

Sep 30

Wednesday

Sep 29

Tuesday

Sep 23

Wednesday

Sep 22

Tuesday

Sep 2

Wednesday
  1. BenchMIRT: What are LLM benchmarks actually measuring?

    Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.

Aug 28

Friday

Aug 27

Thursday
  1. Piloting the world's first double-blind AI evaluations

    Google DeepMind announced the world’s first double-blind evaluation of proprietary frontier AI models, restricting external evaluations to cryptographically isolated environments to prevent models seeing test questions in advance. The pilot partners with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to test Gemini Flash Lite with confidential benchmarks in a privacy-preserving environment.

Aug 24

Monday
  1. Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

    METR research shows AI's effects on scientific discovery are uneven: cybersecurity vulnerability reporting accelerated sharply in 2026 compared with 2025, mathematics accelerated only modestly, and algorithmic progress in AI research itself showed no significant acceleration. A multi-university team proposed SPADE, alternating LLM generation of executable training environments with solving them. Qwen3-30B-A3B averaged 58.3 on the game-environment suite, 8.1 above baseline.

Aug 21

Friday
  1. Measuring benchmark optimization in speech recognition

    Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.

Aug 17

Monday

Aug 13

Thursday
  1. MindTopo reveals VLMs’ spatial reasoning abilities

    Microsoft Research introduced MindTopo to evaluate multimodal models' reasoning and planning over connectivity, enclosure, order, separation and knotting. Current models perform much better at static-image recognition than interactive planning, with failures mainly in planning rather than perception and overall performance far below humans. Image and video generation helps only when preserving structural relations in a single frame; multi-step operations often change topology or violate constraints.

Aug 12

Wednesday
  1. Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

    Google Research's knowledge profiling framework evaluated 13 LLMs on WikiProfile's 2,150 Wikipedia facts. Gemini-3-Pro and GPT-5 encoded 95–98% but failed direct recall of 26–34%, and still failed 11–12% with thinking enabled. It argues frontier factual errors arise more from knowledge accessibility than absence, shifting the bottleneck from acquisition to use.

Aug 4

Tuesday

Jul 27

Monday

Jul 16

Thursday
  1. Newer Models, Same Advantage

    DharmaOCR scored 0.925 on a Portuguese OCR benchmark, ahead of Mistral OCR4's 0.798 and Unlimited-OCR's 0.7587. It specialised through two-stage training: supervised fine-tuning on Portuguese corpora, then DPO to stabilise inference. The author argues that concentrating parameters on one language remains a structural advantage despite emerging architectures.

Jul 15

Wednesday

Jul 6

Monday

Jul 1

Wednesday

Jun 30

Tuesday

Jun 15

Monday

May 16

Saturday

May 4

Monday

Apr 14

Tuesday
  1. Towards developing future-ready skills with generative AI

    Google released Vantage, a research experiment using generative AI simulated conversations to assess future skills such as problem-solving and collaboration in school and university students, with English registration open on Google Labs. An Executive LLM dynamically guides dialogue and an AI Evaluator scores against rubrics. A study with New York University involving 188 US participants aged 18–25 found AI–expert scoring agreement close to agreement between two human experts.

Apr 9

Thursday

Apr 3

Friday

Apr 1

Wednesday

Mar 17

Tuesday
  1. Testing LLMs on superconductivity research questions

    Google and Cornell published a PNAS study asking GPT-4o, Perplexity, Claude 3.5, Gemini Advanced Pro 1.5, NotebookLM and a custom RAG system 67 expert superconductivity questions. Twelve international experts blindly assessed balance, comprehensiveness, concision, evidence, image relevance and qualitative feedback.

Apr 9

Wednesday