Skip to content
Hugging Face Blog·· 1 h ago

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

AI summary

Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system. Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency.

Selection record

AdmittedSum of both 124 ≥ twice the threshold 120

Source tier
Official, first-hand; this tier's threshold is 60
Pre-filter
passed:介绍大模型MoE训练基础设施
Why it was chosen
Gives measured throughput and memory figures for the MoE training stack, so you can judge whether an open approach to large-scale MoE training is viable.

A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.

Source: Hugging Face Blog · huggingface.co