Alexandr Wang releases Meta's Muse Spark 1.3 Max after safety testing

Meta's Max reasoning setting, which powered its strongest Muse Spark 1.3 benchmark results, is now available in Muse Code and the Meta Model API.

By · Published · Updated

Primary source: Alexandr Wang on X

Why it matters

Wang has turned Meta's best Muse Spark 1.3 configuration from a benchmark preview into a product, testing whether rapid model iteration can drive adoption of Meta's agent stack.

Meta releases Muse Spark 1.3 max after a two-day safety hold

Meta chief AI officer Alexandr Wang (@alexandr_wang) released Muse Spark 1.3's Max reasoning setting on September 4th, two days after Meta introduced Muse Spark 1.3 with its most compute-intensive configuration still behind a safety gate. Wang said Meta had completed the additional testing and told users to select "max" in Muse Spark 1.3 or install Muse Code from the command line. Meta's updated release page now says Max is available in Muse Code and the Meta Model API.

Alexandr Wang on X

The short hold offers an early look at how Wang is running Meta's frontier model program: release quickly, keep the highest-compute option gated, then put the executive responsible for the lab directly in front of developers when the gate opens. Muse Spark 1.3 is the fourth version in five months, following the original model in April, version 1.1 in July and version 1.2 in August.

Wang built his career around the operational bottlenecks behind AI. He co-founded Scale AI with Lucy Guo in Y Combinator's Summer 2016 batch, initially offering an API for human data extraction and categorization. Scale AI later expanded into training data, model evaluation and applied AI. Wang studied artificial intelligence at MIT and worked at Quora, Hudson River Trading and Addepar before building Scale AI.

Meta recruited Wang in 2025 alongside a $14.3 billion investment in Scale AI. Scale AI said the transaction valued it above $29 billion. He now serves as Meta's chief AI officer.

That background matters for Muse Spark. Wang spent nearly a decade selling the premise that better data, evaluation and infrastructure determine whether a model works outside a demo. His Meta mandate has pushed him further up the stack, from supplying model builders to controlling the model, coding harness and API through which agents take action.

The benchmark model becomes a product

Meta introduced Muse Spark 1.3 on September 2nd for long-running coding and agentic work. Meta says the model can gather context from conflicting sources, maintain several workflows inside one thread, repair gaps in a plan, ask for clarification and seek confirmation before consequential actions. Those behaviors target the practical failure modes that emerge when an agent works across repositories, documents and software tools for longer than a single exchange.

Meta also says Muse Spark 1.3 used about 20% fewer tool calls and 25% fewer tokens than Muse Spark 1.2 in comparisons conducted by Meta engineers. The figures are internal measurements, and Meta has not tied them to a complete pricing or latency analysis for Max reasoning. They still describe the commercial target clearly: agents that finish work with fewer paid inference steps are easier to deploy repeatedly.

Max matters because Meta used that configuration for many of Muse Spark 1.3's strongest evaluation results. Meta's evaluation methodology says it ran Muse Spark 1.3 at Max reasoning against Claude Opus 5 and GPT-5.6 Sol at their own maximum settings. Meta also cautions that its runs of third-party models were best-effort and may not reflect provider-optimized performance. Some reported scores came from official leaderboards or model providers rather than a single independent test harness.

The gap between Max and the xhigh setting already available on September 2nd varies by task. VentureBeat reported Meta scores of 66.9 for Max and 57.2 for xhigh on OSWorld 2.0, while the two settings tied at 89.4 on DeepSearchQA. On Terminal-Bench 2.1, xhigh narrowly led Max in Meta's results. Max adds capability in places, rather than producing a uniform jump across every benchmark.

Artificial Analysis scored Max at 62 and xhigh at 61 on its Intelligence Index. Its September 2nd release analysis described Max as a limited partner preview. As of September 4th, Artificial Analysis lists Muse Spark 1.3 in both xhigh and Max variants, reflecting the Max configuration's subsequent release.

Wang's two-day safety gate

Meta attributed the delayed Max release to additional safety testing. Wang wrote on X that Meta was launching after completing that work. The sequence gives Wang a practical safety posture to defend: Meta shipped the standard reasoning modes on September 2nd, held back the configuration that spends more compute on difficult tasks, and released it after a separate testing period. Wang told Axios during the initial launch that Meta had "not yet had to pause" its broader work over safety concerns. The Max gate shows Meta is still willing to stagger access at the configuration level.

That distinction will matter as Meta works toward agents with access to personal files, communications and accounts. The release gives Wang a stronger model for that direction while raising the cost of mistakes as the software receives broader operational responsibility.

Meta is selling the whole agent stack

Muse Spark 1.3 also continues Wang's shift away from treating the model weights as the entire product. Meta is distributing this release through Muse Code and the Meta Model API, keeping the model proprietary for now. Meta's roadmap says an open-weights Muse Spark release is planned, without attaching a model version or date.

For Wang, Max is a small release with a large strategic job. Meta needs developers to treat Muse Code and its API as credible alternatives to Claude Code, OpenAI Codex and Cursor, while Meta's consumer products provide a potential distribution channel that independent coding-tool vendors cannot match. Axios reported that Wang sees Muse Spark 1.3 as groundwork for personal agents, tying the developer release to Meta's wider push across its apps and devices.

The September 4th release converts Meta's best Muse Spark 1.3 benchmark configuration into something developers can attempt to use. Wang's next evidence will come from whether teams move sustained coding and agent workloads into Meta's stack. A two-day safety gate is easy to clear. Winning recurring work from developers will take longer.

Reader comments

Conversation for this story loads after sign-in.