Skip to content

All AI news

31 today

Sep 8

Tuesday

Sep 7

Monday

Sep 6

Sunday

Sep 5

Saturday

Sep 4

Friday
  1. Transfer learning for genomic prediction in underrepresented populations

    Google Research evaluated cross-population transfer of polygenic risk scores (PRS) on eight clinical measures using European UK Biobank data and nearly 200,000 Japanese participants from Biobank Japan. At target-population samples of 15,000 or more, target-specific models outperformed mixed European-data training, while highly genetically correlated traits continued to benefit from European data until target samples reached 25,000–40,000 or more.

Sep 3

Thursday
  1. Introducing WeatherNext 3, our most advanced and accurate global weather AI model

    Google DeepMind and Google Research released WeatherNext 3, calling it the most advanced and accurate global weather model available. It learns directly from real-time geostationary Earth-observation satellite data and generates hourly forecasts, with five-kilometre resolution for key surface variables and roughly five times the overall detail of WeatherNext 2.

Sep 2

Wednesday
  1. BenchMIRT: What are LLM benchmarks actually measuring?

    Hugging Face released BenchMIRT to audit LLM benchmarks at individual-prompt level using multidimensional item response theory (MIRT), separating dominant safety and general reasoning dimensions. Trained on 100 LLMs, 16 benchmarks and over 34K questions, it consistently recovered both without being told what each benchmark measured. Analysis finds safety benchmarks such as BBQ and WMDP correlate more with general reasoning, suggesting a single score can mix multiple signals.

Sep 1

Tuesday

Aug 31

Monday

Aug 28

Friday

Aug 27

Thursday
  1. Piloting the world's first double-blind AI evaluations

    Google DeepMind announced the world’s first double-blind evaluation of proprietary frontier AI models, restricting external evaluations to cryptographically isolated environments to prevent models seeing test questions in advance. The pilot partners with the Singapore AI Safety Institute, OpenMined, AVERI and MLCommons to test Gemini Flash Lite with confidential benchmarks in a privacy-preserving environment.

Aug 26

Wednesday
  1. AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR

    Google released research prototype AgentHands, published at CHI 2026, mapping LLM reasoning to speech-synchronised hand animations in XR headsets so agents can point and demonstrate object operations in 3D. It combines environment perception, a gesture event library, gesture-embedding reasoning and local synchronised execution, supporting deictic, iconic and expressive gestures. In an N=12 user study, it significantly improved spatial reference over voice-only interaction.

Aug 25

Tuesday
  1. Mistral x HUMAIN

    Mistral and HUMAIN announced a strategic partnership covering infrastructure, advanced models and deployment, initially focusing on cybersecurity and voice, with frontier models strong in Arabic planned. Worth hundreds of millions of euros, it includes exploring HUMAIN data centres and joint market strategies for regulated Saudi industries.

Aug 24

Monday
  1. Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

    METR research shows AI's effects on scientific discovery are uneven: cybersecurity vulnerability reporting accelerated sharply in 2026 compared with 2025, mathematics accelerated only modestly, and algorithmic progress in AI research itself showed no significant acceleration. A multi-university team proposed SPADE, alternating LLM generation of executable training environments with solving them. Qwen3-30B-A3B averaged 58.3 on the game-environment suite, 8.1 above baseline.

Aug 22

Saturday
  1. An AI tool for prioritizing candidate biomarkers from wearable sensor data

    Google introduced Biomarker Discovery Framework, a human-supervised multi-agent system organising candidate biomarker prioritisation into iterative research cycles. Across 9,279 participant observations in three cohorts, it automatically identified 41 candidate digital mental health biomarkers and 25 metabolic candidates, including an association between sleep-duration variability and PHQ-8 severity (ρ = 0.252).