Skip to content

All AI news

5 today

Sep 2

Wednesday
  1. BenchMIRT: What are LLM benchmarks actually measuring?

    Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.

Sep 1

Tuesday

Aug 28

Friday

Aug 27

Thursday
  1. Piloting the world's first double-blind AI evaluations

    Google DeepMind announced the world’s first double-blind evaluation of proprietary frontier AI models, restricting external evaluations to cryptographically isolated environments to prevent models seeing test questions in advance. The pilot partners with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to test Gemini Flash Lite with confidential benchmarks in a privacy-preserving environment.

Aug 26

Wednesday
  1. AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

    Google released research prototype AgentHands, published at CHI 2026, mapping LLM reasoning to speech-synchronised hand animations in XR headsets so agents can point and demonstrate object operations in 3D. It combines environment perception, a gesture event library, gesture-embedding reasoning and local synchronised execution, supporting deictic, iconic and expressive gestures. In an N=12 user study, it significantly improved spatial reference over voice-only interaction.

Aug 25

Tuesday
  1. Mistral x HUMAIN

    Mistral and HUMAIN announced a strategic partnership covering infrastructure, advanced models and deployment, initially focusing on cybersecurity and voice, with frontier models strong in Arabic planned. Worth hundreds of millions of euros, it includes exploring HUMAIN data centres and joint market strategies for regulated Saudi industries.

Aug 22

Saturday
  1. An AI tool for prioritizing candidate biomarkers from wearable sensor data

    Google introduced Biomarker Discovery Framework, a human-supervised multi-agent system organising candidate biomarker prioritisation into iterative research cycles. Across 9,279 participant observations in three cohorts, it automatically identified 41 candidate digital mental health biomarkers and 25 metabolic candidates, including an association between sleep-duration variability and PHQ-8 severity (ρ = 0.252).

Aug 21

Friday
  1. Measuring benchmark optimization in speech recognition

    Hugging Face proposed three tests to quantify benchmark optimisation in speech recognition: a consensus-disagreement probe, masked-entity retrieval and spelling switches. Evaluating 11 open-source ASR models, it found some high-scoring models reproduce errors in VoxPopuli and LibriSpeech reference transcripts even when contradicted by audio, when relevant words are muted or when both spellings fit the audio.

Aug 20

Thursday

Aug 19

Wednesday

Aug 18

Tuesday

Aug 17

Monday

Aug 14

Friday
  1. State of Open Models: Summer 2026 Observations

    Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.

Aug 13

Thursday
  1. MindTopo reveals VLMs’ spatial reasoning abilities

    Microsoft Research introduced MindTopo to evaluate multimodal models' reasoning and planning over connectivity, enclosure, order, separation and knotting. Current models perform much better at static-image recognition than interactive planning, with failures mainly in planning rather than perception and overall performance far below humans. Image and video generation helps only when preserving structural relations in a single frame; multi-step operations often change topology or violate constraints.

Aug 12

Wednesday
  1. Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

    Google Research's knowledge profiling framework evaluated 13 LLMs on WikiProfile's 2,150 Wikipedia facts. Gemini-3-Pro and GPT-5 encoded 95–98% but failed direct recall of 26–34%, and still failed 11–12% with thinking enabled. It argues frontier factual errors arise more from knowledge accessibility than absence, shifting the bottleneck from acquisition to use.

  2. Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    Microsoft Research released CARE-X, a unified chest X-ray vision-language research model for report generation and structured prediction. It rewards clinical correctness using multitask reinforcement learning (DAPO). Generation and dual-inference modes cover lesion presence and negation, localisation, multilabel classification, catheter and tube malposition detection, and localisation of 29 anatomical regions.