Skip to content

Data and training

The training side: building datasets, synthetic data, pre- and post-training methods, compute and training cost.

30selectedRelated topicsResearchDeploymentModel releases

Latest selected

1–20 of 30

Sep 30

Wednesday

Sep 29

Tuesday
  1. NVIDIA Kumo Tabular Sets a New Accuracy-Efficiency Frontier for Tabular Prediction

    NVIDIA released Kumo Tabular, an open-source tabular foundation model that predicts labels for new rows in a single forward pass given labelled rows, without training, tuning or feature engineering. It supports classification and regression, offers three sizes from 28M to 215M, and was pretrained solely on artificially generated tables. It uses the commercially usable OpenMDW-1.1 licence and ranks first on TabArena, BeyondArena, TALENT and ScoringBench.

Sep 24

Thursday

Sep 22

Tuesday

Sep 16

Wednesday

Sep 8

Tuesday

Sep 4

Friday

Aug 28

Friday

Aug 26

Wednesday

Aug 25

Tuesday

Aug 19

Wednesday

Aug 14

Friday
  1. State of Open Models: Summer 2026 Observations

    Hugging Face published its open-source model observatory report for January–August 2026. Hub data shows Chinese labs released the largest open-source models by parameter count in most months, with Chinese monthly peaks between 754 billion and 2.78 trillion parameters, while US models stayed below 130 billion in five of seven months.

Aug 6

Thursday

Jul 31

Friday

Jul 9

Thursday

Jun 30

Tuesday