Skip to content

#Data and training

1 today

Aug 24

Monday
  1. Import AI 470: No rights for machines; automating environment generation with SPADE; and building better GPU kernels with Hawkeye

    METR research shows AI's effects on scientific discovery are uneven: cybersecurity vulnerability reporting accelerated sharply in 2026 compared with 2025, mathematics accelerated only modestly, and algorithmic progress in AI research itself showed no significant acceleration. A multi-university team proposed SPADE, alternating LLM generation of executable training environments with solving them. Qwen3-30B-A3B averaged 58.3 on the game-environment suite, 8.1 above baseline.

Aug 22

Saturday
  1. An AI tool for prioritizing candidate biomarkers from wearable sensor data

    Google introduced Biomarker Discovery Framework, a human-supervised multi-agent system organising candidate biomarker prioritisation into iterative research cycles. Across 9,279 participant observations in three cohorts, it automatically identified 41 candidate digital mental health biomarkers and 25 metabolic candidates, including an association between sleep-duration variability and PHQ-8 severity (ρ = 0.252).

Aug 21

Friday

Aug 19

Wednesday

Aug 18

Tuesday

Aug 17

Monday

Aug 14

Friday
  1. State of Open Models: Summer 2026 Observations

    Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.

Aug 13

Thursday

Aug 12

Wednesday
  1. Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

    Google Research's knowledge profiling framework evaluated 13 LLMs on WikiProfile's 2,150 Wikipedia facts. Gemini-3-Pro and GPT-5 encoded 95–98% but failed direct recall of 26–34%, and still failed 11–12% with thinking enabled. It argues frontier factual errors arise more from knowledge accessibility than absence, shifting the bottleneck from acquisition to use.

  2. Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    Microsoft Research released CARE-X, a unified chest X-ray vision-language research model for report generation and structured prediction. It rewards clinical correctness using multitask reinforcement learning (DAPO). Generation and dual-inference modes cover lesion presence and negation, localisation, multilabel classification, catheter and tube malposition detection, and localisation of 29 anatomical regions.

Aug 10

Monday

Aug 6

Thursday

Jul 31

Friday

Jul 30

Thursday

Jul 26

Sunday

Jul 22

Wednesday

Jul 16

Thursday
  1. Newer Models, Same Advantage

    DharmaOCR scored 0.925 on a Portuguese OCR benchmark, ahead of Mistral OCR4's 0.798 and Unlimited-OCR's 0.7587. It specialised through two-stage training: supervised fine-tuning on Portuguese corpora, then DPO to stabilise inference. The author argues that concentrating parameters on one language remains a structural advantage despite emerging architectures.

Jul 9

Thursday

Jul 7

Tuesday

Jul 6

Monday
  1. PRX Part 4: Our Data Strategy

    Hugging Face's fourth PRX instalment details its data strategy: mixing public and internal pretraining datasets, regenerating long image captions with a VLM and converting them for streaming training. It uses Lance for construction and filtering and MDS for streaming. After switching to Qwen3-VL, text latents are computed during training, with measured throughput loss of around 3–4%, or roughly one extra day for 30 days of training.

Jul 1

Wednesday

Jun 30

Tuesday

Jun 17

Wednesday
  1. From pixels to planning: Earth AI for nature restoration

    Google Research released vector data converting Farmscapes 2020 high-resolution raster maps into usable inventories of hedgerows, stone walls and coppices across over 130,000 square kilometres of the UK. The framework fine-tunes an RSF Vision-Transformer pretrained on over 300 million global satellite images with around 247 square kilometres of labelled data. Polsby–Popper compactness distinguishes woodland, clusters of trees and hedgerows, with a threshold below 0.5 for linear features.

Jun 8

Monday

Jun 4

Thursday
  1. The next chapter in flood resilience: Open sourcing Google’s hydrology framework

    Google Research open-sourced its hydrological modelling framework on GitHub under Apache 2.0, enabling national weather and hydrology agencies to integrate AI flood forecasting. The Python package uses PyTorch and an LSTM architecture, can train or fine-tune on Caravan data, and includes interactive tutorial notebooks and videos.

May 27

Wednesday
  1. Introducing physics AI at Mistral: the foundation for engineering acceleration.

    After incorporating Emmi AI, Mistral launched physical AI capabilities for AI-native industrial engineering with partners including ASML, Airbus, Safran and Siemens Energy. The model predicts physical fields directly from geometry and boundary conditions in seconds through one forward pass on a single GPU, versus hours to weeks per design variant in traditional CFD/FEM. It accelerates design iteration while retaining traditional solvers for validation and edge cases.

May 23

Saturday
  1. Emmi joins Mistral to accelerate the AI-native industry

    Mistral AI signed a definitive agreement this week to acquire Physics AI pioneer Emmi AI, strengthening its AI transformation services for industrial companies. Founded in Austria, Emmi AI has over 30 researchers and engineers focusing on large engineering models that replace days of computation with real-time simulation and build digital twins. Its co-founders and team will join Mistral's Science and Applied AI teams in May.

May 20

Wednesday

May 18

Monday

May 16

Saturday
  1. Finding the molecular switches behind new infectious diseases

    Cambridge professor Clare Bryant uses Google Co-Scientist to study molecular mechanisms causing sepsis when pathogens such as influenza cross species. Generated and ranked hypotheses identified a previously overlooked protein and then specific amino acid sites. Her team is building cell lines carrying these mutations to test the hypotheses, expecting work that normally takes two to three years to finish in six months.

May 4

Monday

May 2

Saturday
  1. Catalyzing scientific impact through global partnerships and open resources

    Google Research said its open-source tools and datasets have enabled over 250,000 researchers and developers worldwide. Genomics tools DeepVariant, DeepConsensus and DeepPolisher have supported processing exomes and whole genomes from 2.5 million people; MedGemma has over 4.8 million downloads. Open Health Stack is deployed in more than ten countries, reaching over 65 million beneficiaries.

Apr 30

Thursday