Allen Institute releases Olmo-core 3, an MoE training stack that scales to a trillion parameters

without collapsing The Allen Institute for AI released Olmo-core 3 today, the third version of the development framework behind its Olmo models. At its center is an open mixture-of-experts training system designed to reach the trillion-parameter range while preserving computational efficiency.
The bottleneck is well known: MoE architectures let developers grow parameter capacity without activating every expert for each token, but the full model still has to reside in GPU memory and be updated during training. Routing inputs to the right experts across a compute cluster creates communication overhead that erodes the theoretical advantage. Olmo-core 3 was built to close that gap.
In an internal benchmark, researchers scaled the expert pool from 8 to 128 while selecting only four experts per token — the smallest text units the model processes — keeping active parameters per token stable at roughly 3.2 billion. Total capacity jumped from 4.6 billion to 47 billion parameters, while training throughput fell less than five percent. The same infrastructure has already been tested at more than one trillion total parameters, signaling a clear path forward.
The core architectural change is a shift from fully sharded data parallelism (FSDP), which gathered and redistributed weights for every micro-batch, to distributed data parallelism (DDP), which keeps experts resident on the GPU and streams only the relevant data to them. Against the established commercial alternative, Nvidia's Megatron-Core, Olmo-core 3 offers an integrated stack inside the Olmo framework. In a preliminary run on eight Nvidia B300 GPUs, a 47-billion-parameter MoE model processed 52 thousand tokens per second per GPU with the new stack, versus 19,400 with the previous version — a 2.7× throughput gain.
The new stack combines three parallelism techniques to split the model and training state across a GPU cluster: expert parallelism, which distributes experts across cards so each GPU holds only a slice of the pool; pipeline parallelism, which divides the model's layers — the sequential stages that process the input — across GPU groups and reduces the memory each card needs; and a distributed optimizer, which shards the optimizer state — the auxiliary data used to compute and apply updates — instead of keeping a full copy on every GPU. Together they let MoE models grow without requiring every GPU to hold the full model and training state in memory.
The release follows two earlier generations: OlmoE used an MoE architecture with 64 routed experts, while Olmo 3 moved to a dense architecture in which nearly the entire model is active for every token, and its stack was built accordingly. Olmo-core 3 returns the focus to sparse models at a far larger scale, part of the institute's stated commitment to open the tools and infrastructure behind each new generation. For academic researchers and smaller labs that have been pushed out of advanced model development by compute and energy costs, it is infrastructure that lets them experiment with distributed training on the hardware they already have, without reinventing the wheel.