Skip to content

#Data and training

1 today

Oct 1

ThursdayToday1 items

Sep 30

Wednesday

Sep 29

Tuesday
  1. NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction

    NVIDIA released Kumo Tabular, an open-source tabular foundation model that predicts labels for new rows in a single forward pass given labelled rows, without training, tuning or feature engineering. It supports classification and regression, offers three sizes from 28M to 215M, and was pretrained solely on artificially generated tables. It uses the commercially usable OpenMDW-1.1 licence and ranks first on TabArena, BeyondArena, TALENT and ScoringBench.

Sep 28

Monday

Sep 26

Saturday

Sep 25

Friday

Sep 24

Thursday

Sep 23

Wednesday
  1. Sakeena Fiza Helps NVIDIA Hardware Succeed at Scale

    NVIDIA validation engineer Sakeena Fiza validates new hardware before mass production in its data centre systems lab. One memorable moment was the first successful system-level enumeration of the Rubin GPU. She compares validation to solving crimes, finding and reproducing issues before customers do, across trays, racks, clusters and customers' AI factories.

Sep 22

Tuesday

Sep 21

Monday

Sep 16

Wednesday

Sep 11

Friday

Sep 10

Thursday

Sep 8

Tuesday

Sep 4

Friday
  1. Transfer learning for genomic prediction in underrepresented populations

    Google Research evaluated cross-population transfer of polygenic risk scores (PRS) on eight clinical measures using European UK Biobank data and nearly 200,000 Japanese participants from Biobank Japan. At target-population samples of 15,000 or more, target-specific models outperformed mixed European-data training, while highly genetically correlated traits continued to benefit from European data until target samples reached 25,000–40,000 or more.

Sep 3

Thursday

Sep 2

Wednesday

Sep 1

Tuesday

Aug 28

Friday

Aug 27

Thursday

Aug 26

Wednesday

Aug 25

Tuesday

Aug 22

Saturday
  1. An AI tool for prioritizing candidate biomarkers from wearable sensor data

    Google introduced Biomarker Discovery Framework, a human-supervised multi-agent system organising candidate biomarker prioritisation into iterative research cycles. Across 9,279 participant observations in three cohorts, it automatically identified 41 candidate digital mental health biomarkers and 25 metabolic candidates, including an association between sleep-duration variability and PHQ-8 severity (ρ = 0.252).

Aug 21

Friday

Aug 19

Wednesday

Aug 18

Tuesday

Aug 14

Friday
  1. State of Open Models: Summer 2026 Observations

    Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.

Aug 13

Thursday

Aug 12

Wednesday
  1. Empty shelves or lost keys? Recall is the bottleneck for parametric factuality

    Google Research's knowledge profiling framework evaluated 13 LLMs on WikiProfile's 2,150 Wikipedia facts. Gemini-3-Pro and GPT-5 encoded 95–98% but failed direct recall of 26–34%, and still failed 11–12% with thinking enabled. It argues frontier factual errors arise more from knowledge accessibility than absence, shifting the bottleneck from acquisition to use.

  2. Introducing CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    Microsoft Research released CARE-X, a unified chest X-ray vision-language research model for report generation and structured prediction. It rewards clinical correctness using multitask reinforcement learning (DAPO). Generation and dual-inference modes cover lesion presence and negation, localisation, multilabel classification, catheter and tube malposition detection, and localisation of 29 anatomical regions.