Skip to content
Hugging Face Blog·· 2026-07-06

PRX Part 4: Our Data Strategy

PRX Part 4: Our Data Strategy

AI summary

Hugging Face's fourth PRX instalment details its data strategy: mixing public and internal pretraining datasets, regenerating long image captions with a VLM and converting them for streaming training. It uses Lance for construction and filtering and MDS for streaming. After switching to Qwen3-VL, text latents are computed during training, with measured throughput loss of around 3–4%, or roughly one extra day for 30 days of training.

Selection record

Not admittedSum of both 68 < twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:详述AI模型PRX数据策略与训练

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Hugging Face Blog · huggingface.co