Skip to content
Hugging Face Blog·· 2026-09-02

BenchMIRT: What are LLM benchmarks actually measuring?

BenchMIRT: What are LLM benchmarks actually measuring?

AI summary

Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.

Selection record

Not admittedSum of both 65 < twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:介绍LLM基准评测新方法BenchMIRT

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Hugging Face Blog · huggingface.co