SpaceXAI says Grok 4.8 will start reinforcement learning this week

Elon Musk says the 2.5T run beat a 2.1T JAX-trained model despite mistakes corrected midway; a 3T successor is already planned.

By · Published

Primary source: X - Elon Musk

Why it matters

SpaceXAI is treating training software and data quality as scaling levers alongside model size. Musk's mid-run corrections leave the performance claim unproven until evaluations appear.

An abstract, glowing visualization of data streams and patterns representing an artificial intelligence model undergoing reinforcement learning in a high-tech environment.

Elon Musk (@elonmusk) said in a two-post thread on X on September 14th that Grok 4.8 will finish its current training stage during the week of September 14th and move into reinforcement learning, marking the next step toward a release of SpaceXAI's new frontier model.

Musk called Grok 4.8 a "2.5T model" trained with a new internal C++ software stack. He said it was "unequivocally better" than an earlier 2.1T run trained with JAX, though he provided no evaluations or test results supporting that comparison. Musk also acknowledged that SpaceXAI made mistakes during the 2.5T run and corrected them only midway through training.

Musk co-founded xAI in 2023 and was still identified by xAI as its co-founder and CEO in a September 2025 announcement. The AI operation became part of SpaceX in a February 2nd, 2026 acquisition, and the combined organization now operates as SpaceXAI.

The software stack is the news

The move from JAX to an internal C++ stack is the most consequential detail in Musk's update. JAX is a Python library for large-scale numerical computing and machine learning that compiles workloads for GPUs and other accelerators. It gives researchers a high-level programming model while handling differentiation, compilation and parallel execution.

A purpose-built C++ system moves more of that machinery under SpaceXAI's control. At the scale Musk described, training software determines how efficiently accelerators communicate, how quickly engineers recover from failures and how much expensive computing time produces usable model progress. Musk did not publish throughput, reliability or compute-efficiency figures for the new stack, so its advantage over SpaceXAI's JAX implementation cannot yet be measured externally.

The "2.5T" label also requires caution. Musk did not define what the figure measures. Total parameters, active parameters, training tokens and computing operations describe different parts of a model run, and the shorthand alone does not establish Grok 4.8's architecture or its cost. The comparison with a 2.1T run is therefore meaningful primarily as SpaceXAI's internal assessment.

Musk's admission of mid-run mistakes complicates that comparison further. A checkpoint produced after engineers altered the training setup is not a controlled test of C++ against JAX. Data changes add another variable: Musk said SpaceXAI has "greatly improved" data quality over the past few months, making it impossible to isolate how much of the claimed gain came from model scale, software, corrected errors or the training mixture.

Grok 4.8 is still in the pipeline

SpaceXAI released Grok 4.6 on August 12th, making it the current public flagship listed in SpaceXAI's model documentation. SpaceXAI said Grok 4.6 went through supplemental training, supervised fine-tuning and reinforcement learning before release, with its RL work focused on coding, knowledge work and specialized technical environments.

Grok 4.8 has yet to begin that reinforcement-learning stage. Reinforcement learning is where SpaceXAI can train the model against reward signals and task outcomes, including reasoning, coding and agentic workloads. The length of that work, the subsequent safety evaluations and the deployment process will determine when Grok 4.8 reaches users. Musk committed only to starting RL during the week of September 14th.

SpaceXAI is already preparing a larger run. Musk said an upcoming 3T model will use further-improved internal training software and what he described as substantially better data. He did not attach a Grok version number to that run.

The sequence puts SpaceXAI's strategy in plain view: increase scale, replace general-purpose training infrastructure with internal systems, improve the data mixture and carry fixes directly into the next run. Grok 4.8 will be the first public test of whether that C++ rewrite produces a better model rather than simply a larger internal engineering project. That judgment will require reproducible evaluations and deployment details after reinforcement learning is complete.

Reader comments

Conversation for this story loads after sign-in.