Ai2 opens MoE training stack after testing a 1.2T-parameter setup

Ai2 says its system can train larger sparse models with less overhead. Its trillion-parameter result measures system performance, not a completed model.

By · Published

Primary source: Hugging Face Newsroom

Why it matters

Ai2's release makes MoE training infrastructure part of its open-model strategy, but the trillion-parameter headline is a system benchmark, not a trained model. The distinction sets a useful bar for judging both its technical claims and the access the code actually provides.

A generic compute rack with grouped server trays and fiber links represents Ai2’s newly opened MoE training stack.

The Allen Institute for AI (Ai2) released Olmo-core 3 on October 1st, opening a redesigned training system for mixture-of-experts models and reporting benchmarks up to 1.2 trillion total parameters. The result is a systems test, not evidence that Ai2 has trained a trillion-parameter model to useful quality.

Ai2 was founded in 2014 by Microsoft co-founder Paul Allen as a Seattle-based nonprofit research institute. Olmo-core 3 extends Allen's founding commitment to openly shared AI research. Ai2 says the new framework will underpin the next generation of Olmo, which it plans to build as a mixture-of-experts model. In a Hugging Face announcement, Ai2 presents the framework, technical report and interactive walkthrough alongside the public code.

The release targets a practical constraint in sparse models. A mixture-of-experts, or MoE, routes each token to a subset of a much larger pool of specialized model components. That can keep the amount of computation per token lower than a dense model of similar total capacity. But training still requires keeping model weights and optimizer state in memory, while moving tokens to the right experts across GPUs adds communication overhead. Those costs can eat into the savings from activating only part of the model.

Ai2 says its new stack addresses that overhead by changing how the system distributes work. Its earlier MoE implementation used fully sharded data parallelism, gathering and resharing weights for each small batch. Olmo-core 3 instead uses a distributed data parallel approach that keeps experts on GPUs and routes data to them. Ai2 combines that design with expert and pipeline parallelism, a distributed optimizer, GPU-resident routing and grouped matrix computations.

Diagram comparing Ai2's earlier fully sharded MoE approach with Olmo-core 3's distributed data parallel approach, alongside the additional techniques Ai2 says it combines.
Ai2 says Olmo-core 3 keeps experts on GPUs and routes data to them, replacing its earlier approach of gathering and resharing weights for each small batch - AI explanatory diagram, not documentary evidence. RuntimeWire - AI-generated diagram.

The benchmark is the claim

In one comparison, Ai2 increased the expert pool from eight to 128 while selecting four experts for each token. Ai2 says active parameters stayed near 3.2 billion as total capacity grew from 4.6 billion to 47 billion, with training throughput falling by less than 5%.

A separate preliminary test compared the new system with Ai2's earlier implementation on eight NVIDIA B300 GPUs. For a 47-billion-parameter MoE, Ai2 reports 52,000 tokens per second per GPU, against 19,400 with the previous stack, or about 2.7 times the throughput. A four-GPU test of MXFP8, a lower-precision number format, showed about 21% higher throughput than the BF16 baseline and peak active memory declining from 103 GiB to 95 GiB. These are Ai2's own results; the announcement does not establish independent reproduction or disclose the benchmarks' compute cost, energy use or duration.

Table of Ai2-reported preliminary benchmark results: throughput for a 47-billion-parameter MoE on eight NVIDIA B300 GPUs, and throughput and peak active memory in a separate four-GPU MXFP8 test.
Ai2 reports these results from separate preliminary tests; the announcement does not establish independent reproduction - AI explanatory infographic, not documentary evidence. RuntimeWire - AI-generated infographic.

The largest configuration needs careful reading. Ai2 reports testing a 1.2-trillion-parameter setup across 512 B300 GPUs, with 58.36 billion parameters active per token and a peak observed throughput of 858 TFLOP/s per GPU. Ai2 says it used random routing to measure system performance, rather than model quality. A separate short-capacity test reached 2.38 trillion parameters using DeepEP v2, which the institute explicitly describes as a configuration test rather than a full training run. Neither figure should be read as a trained, evaluated model available to use.

Scaling expert count while keeping active computation roughly stable tests whether routing and communication can keep pace as a model's total capacity expands. In Ai2's framing, infrastructure is valuable because a model's architecture is only part of the cost: data movement, memory and coordination determine whether sparse computation pays off in practice.

An open stack for the next Olmo

Olmo-core 3 follows earlier Ai2 work on sparse models, including OlmoE, which used 64 routed experts. The institute says the more recent Olmo 3 model used a dense architecture; its next Olmo generation is planned as an MoE. The new training system therefore serves both as a public research tool and as groundwork for Ai2's own next model.

The project extends Allen's founding premise into a layer that often remains harder to inspect than model weights. Ai2 argues that open model development is more useful when researchers can inspect the training infrastructure and decisions behind a model alongside its final parameters. The Olmo-core repository makes the framework available to researchers and developers, and the interactive walkthrough illustrates how data, experts and model layers are split across GPUs.

Code alone does not remove the hardware barrier for smaller labs. Ai2's largest reported benchmark used 512 high-end GPUs, and the announcement provides no cost figure that would establish how accessible a comparable run is in practice. The nearer-term contribution is a more inspectable implementation and a set of reported performance results that other teams can scrutinize and attempt to reproduce. For Ai2, the immediate payoff is also internal: it has released the infrastructure it says will support its next Olmo model before that model arrives.

Reader comments

Conversation for this story loads after sign-in.