Ai2 Releases Olmo-core 3, Open Training Stack for Trillion-Parameter MoEs
The Allen Institute for AI released Olmo-core 3, an open mixture-of-experts training stack benchmarked at over one trillion total parameters. On 512 NVIDIA B300 GPUs, a 1.2-trillion-parameter model with 58.36 billion active parameters reached 858 TFLOP/s per GPU, while an 8-GPU test showed 2.7x higher throughput than the earlier implementation.
- Benchmark: 1.2T parameters, 58.36B active per token, 512 B300 GPUs
- On 8 B300 GPUs: 52,000 tokens/s per GPU vs 19,400 previously
- MXFP8 boosts throughput 21% and cuts peak memory from 103 to 95 GiB
- Expert pool grew from 8 to 128 with 4 active per token, throughput down under 5%
Read next
AI