Four Grok 4.5 checkpoints appear in Arena near xAI's 4.6 window

The roster entries point to multimodal and web-enabled variants, but their labels do not establish that Grok 4.6 is behind them.

By · Published

Primary source: X - leo (@synthwavedd)

Why it matters

Multiple checkpoints suggest xAI is using live user preferences to choose its next Grok configuration, but the shared 4.5 label is not proof of a 4.6 launch.

Four Grok 4.5 checkpoints appear in Arena near xAI's 4.6 window — The roster entries point to multimodal and web-enabled variants, but their labels do not establish that Grok 4.6 is behind them.

Four checkpoints labeled Grok 4.5 appeared in an Arena roster alert on Tuesday, giving Elon Musk (@elonmusk) a fresh pool of real-user tests as xAI prepares the next iteration of Grok.

The four entries were surfaced in a post on X by leo (@synthwavedd), who described them as presumed Grok 4.6 variants running under the existing model name. The attached roster-monitoring screenshot supports the narrower claim: four separate identifiers were added under "grok-4.5." It does not identify them as Grok 4.6. (x.com)

A roster alert listing four Grok 4.5 checkpoints
A RAGtrack alert lists four Arena roster entries carrying the Grok 4.5 name. Screenshot: leo on X.

Two of the entries are marked for file, image and text inputs with text and web outputs. The other two accept image and text inputs and return text. Each carries a different internal identifier, indicating four distinct endpoints or configurations rather than duplicate listings.

Musk identified a larger model training run as Grok 4.6 in a July 17th reply on X. The appearance of multiple checkpoints less than four weeks later fits a late-stage evaluation cycle, when a lab can compare system prompts, post-training recipes, tool access or different snapshots before selecting a release candidate. The roster data alone cannot determine which of those variables separates the four entries.

Arena lets labs test several candidates

Arena's evaluation policy explicitly allows model providers to test multiple pre-release variants. Unreleased models can be given anonymous labels, evaluated until enough votes accumulate, and removed after Arena shares the results privately with the provider. Arena says each anonymous candidate receives a unique label, while at least one model in every battle must already be publicly available. (arena.ai)

That process turns Arena into an external product-selection system for frontier labs. Users submit their own prompts, compare two responses without initially seeing the model names and vote for the stronger answer. The identities are revealed after the vote. Arena then aggregates those comparisons into ratings, giving labs feedback from live coding and chat workloads that internal benchmark suites may miss. (help.arena.ai)

The four Grok endpoints therefore show active experimentation. They do not prove that xAI has finalized Grok 4.6, settled its specifications or begun a public rollout. The shared Grok 4.5 label could cover release candidates for a successor, updated 4.5 configurations or a mix of both.

Grok 4.5 set the baseline in July

SpaceXAI, the branding used on xAI's current model pages, released Grok 4.5 on July 16th, positioning it around coding, agentic work and technical knowledge. SpaceXAI says Grok 4.5 was trained across tens of thousands of Nvidia GB300 GPUs and subjected to reinforcement learning on hundreds of thousands of tasks. SpaceXAI priced the API at $2 per million input tokens and $6 per million output tokens. Those training, performance and efficiency figures remain company-supplied claims. (x.ai)

Grok 4.5 was already established inside Arena before Tuesday's roster alert. Arena's public changelog recorded the model's addition to Agent Arena on July 13th, three days before SpaceXAI's formal launch announcement. The sequence provides a recent precedent for Grok appearing in Arena ahead of, or alongside, a broader release. (arena.ai)

Arena's late-July WebDev leaderboard placed the public Grok 4.5 endpoint 10th overall, with a score of 1,550 and a confidence range spanning seventh through 16th place. The model's four new checkpoints give xAI a way to test whether its next configuration can close that gap on coding tasks while preserving the speed and pricing that SpaceXAI emphasized at launch. (arena.ai)

For developers, the actionable detail is limited to the roster change. There is no verified Grok 4.6 model card, API identifier, price or public endpoint attached to these four entries. What has surfaced is the evaluation phase: xAI appears to be testing several Grok configurations across general chat, multimodal input, file handling and web-enabled output before deciding what ships under the next name.

Reader comments

Conversation for this story loads after sign-in.