News

Introducing Olmo-core 3: Open, Scalable Training Infrastructure for Large MoEs

AllenAI, in collaboration with Hugging Face, has introduced Olmo-core 3, an open and scalable training infrastructure purpose-built for large Mixture of Experts (MoE) models. Mixture of Experts architectures have gained prominence for their ability to scale model capacity efficiently by activating only subset of model parameters during inference, enabling larger effective model sizes without proportional computational costs.

The release emphasizes openness, allowing researchers and developers to access and build upon the training infrastructure. Scalability remains a core design principle, supporting the computational demands of training cutting-edge MoE systems.

This development aligns with broader industry efforts to make advanced AI training methodologies more accessible to the research community, potentially lowering barriers for institutions seeking to experiment with large-scale MoE architectures.

Further technical details regarding specifications, performance benchmarks, and availability are expected in the official documentation on Hugging Face.

Sources