SpaceXAI ships Grok 4.6 with a 500,000-token context window

The model keeps Grok 4.5's $2/$6 base API price, adds a 500,000-token context window and launches in Cursor and Grok Build.

By · Published

Primary source: VentureBeat

Why it matters

SpaceXAI is pairing Grok 4.6 with its Grok Build agent environment and Cursor distribution while SpaceX's proposed Anysphere acquisition remains pending. Its larger context window, benchmark gains and long-context price increase give developers concrete factors to assess against production reliability and total workload cost.

An advanced AI model (Grok 4.6) being deployed and integrated into various platforms, powering numerous autonomous agents across diverse digital tasks. (Museum miniature diorama, handcrafted from painted wood, paper, and sculpted clay, meti

Elon Musk's SpaceXAI (on X) announced Grok 4.6 on Wednesday, August 12, extending its rapid model-release schedule with a system designed to stay on task through longer coding and knowledge-work jobs.

The August 12th announcement positions Grok 4.6 as a practical agent model: one that can research unfamiliar subjects, navigate a codebase, operate tools and turn a broad product idea into working software over a sequence of steps. Its base API price remains $2 per million input tokens and $6 per million output tokens, the same rates SpaceXAI set for Grok 4.5.

Those benchmark gains, unchanged base pricing and distribution through Cursor give developers several reasons to evaluate Grok 4.6 for longer coding and knowledge-work jobs. The performance increase gives engineering organizations a reason to rerun model evaluations, while the unchanged entry price allows testing on larger workloads.

VentureBeat reported that Grok 4.6 scored 61 on the Artificial Analysis Intelligence Index, five points above Grok 4.5. The score moved Grok ahead of Moonshot AI's Kimi K3 and level with OpenAI's GPT-5.6 Sol Max, according to the report. SpaceXAI's published evaluation table shows Grok 4.6 at 61 on the Artificial Analysis Intelligence Index and details its benchmark results.

SpaceXAI's announcement confirms that Grok 4.6 was available in Cursor and Grok Build at launch. The release separately lists the API, OpenRouter, Vercel and Cloudflare as distribution channels, without establishing first-day availability across every listed partner.

A founder consolidating the stack

SpaceX announced the xAI acquisition on February 2, 2026. xAI announced a $20 billion Series E in January 2026, with participation from Valor Equity Partners, StepStone Group, Fidelity, Qatar Investment Authority, MGX, Baron Capital, NVIDIA and Cisco Investments.

Musk's operating history spans software, payments, electric vehicles and rockets. Tesla's corporate biography lists him as a University of Pennsylvania physics and business graduate who co-founded Zip2 and PayPal before leading SpaceX and Tesla and founding Neuralink and The Boring Company.

The February merger placed SpaceXAI within the broader SpaceX corporate structure. SpaceXAI offers Grok 4.6 through its API, while Grok Build provides the agent environment. Cursor provides access to software developers, though SpaceX's proposed acquisition of its parent company remains pending.

AP reported that SpaceX agreed to acquire Cursor parent Anysphere for approximately $60 billion in stock. The transaction is expected to close in the third quarter of 2026, subject to closing conditions and regulatory approvals, so Grok 4.6's availability in Cursor predates any completed acquisition.

Cursor supports models from multiple frontier labs, including companies that compete directly with SpaceXAI. Grok therefore competes for usage inside a coding environment that also gives developers access to rival models.

The price stays low until the context gets long

SpaceXAI's headline API rates need one qualification. The $2 input and $6 output prices apply while prompts remain below 200,000 tokens. At or above that threshold, rates double to $4 per million input tokens and $12 per million output tokens. Once a prompt crosses the threshold, the higher rates apply to all tokens in the request.

The model documentation lists a 500,000-token context window, text and image inputs, structured outputs, reasoning and function calling.

For an agent that repeatedly reads a repository, tool results and its own prior work, the 200,000-token boundary can materially change the bill. SpaceXAI is still pricing Grok 4.6 below several premium frontier models in VentureBeat's comparison, but buyers will need to test full trajectories rather than compare the first line of a pricing table.

SpaceXAI is also offering a faster Grok 4.6 variant at twice the standard price. That creates a second commercial option for workloads where lower latency justifies a higher token bill.

Benchmark gains leave a terminal-work gap

SpaceXAI says Grok 4.6 received a longer supplemental training run than Grok 4.5. SpaceXAI used curated model-generated reasoning and technical data, engineering data, an updated optimizer and a revised training recipe. Grok 4.5 then generated new supervised fine-tuning trajectories across reasoning, software engineering, STEM and knowledge work, with model-based checks filtering problematic traces.

Reinforcement learning focused on coding, web development, computer-aided design, kernel optimization and general knowledge work. SpaceXAI says the resulting model performs more self-testing and verification during long jobs. Those behavior claims come from internal testing and require validation in live deployments.

SpaceXAI's published evaluation table shows a substantial generational improvement without a clean sweep. Grok 4.6 scored 65.9% on DeepSWE v1.1, up from 54% for Grok 4.5, while GPT-5.6 Sol Max scored 73%. It reached 69.9% on CursorBench v3.2, compared with 70.5% for Fable 5 Max. On APEX-Agents, Grok 4.6 rose 10.4 percentage points to 57.5%, beating the listed GPT-5.6 Sol Max result of 56.7% and trailing Fable 5 Max at 59.2%.

Terminal-Bench v3.0 remains a visible weakness. Grok 4.6 improved to 26% from 15.7%, while the comparison scores for GPT-5.6 Sol Max and Fable 5 Max were 34.6% and 34.1%, respectively. The model has become more capable at terminal work without closing the gap.

Artificial Analysis describes its capability indices as composites of independently run benchmarks and cautions that any index may fail to map directly onto a particular production use case. A score of 61 is useful evidence for a trial. It does not establish reliability across an organization's repository, tools, permissions and review process.

Grok 4.6 follows Grok 4.5 by four weeks

Grok 4.6 follows Grok 4.5, released on July 16th, and keeps its predecessor's base token pricing. RuntimeWire previously described Grok 4.5 as Musk's price weapon against premium coding models. The new release raises benchmark scores while holding base pricing.

The launch announcement confirms Grok 4.6 availability in Cursor and Grok Build, with twice the included usage in both products for the first week. Its get-started section also identifies the API, OpenRouter, Vercel and Cloudflare as distribution channels. Grok Build access starts with the $30-per-month SuperGrok plan.

Agent performance is shaped by the software around the model: how context is assembled, when tools are called, how failures are detected and what a human must approve. Grok Build gives SpaceXAI a product layer where it can tune that agent experience around Grok 4.6.

SpaceXAI now trains the model, offers the Grok Build agent environment, sells inference and distributes Grok inside Cursor while SpaceX pursues its acquisition of Cursor parent Anysphere. Grok 4.6's benchmark results support further evaluation, but sustained usage will depend on production reliability, total workload cost and performance inside each coding agent.

Reader comments

Conversation for this story loads after sign-in.