Meta says Muse Spark 1.2 gains 12.2 points when given tools

Alexandr Wang's MSL reports that its model trails its predecessor without tools and pulls ahead when paired with Meta's runtime.

By · Published

Primary source: AI at Meta

Why it matters

Meta's chart shows Muse Spark 1.2 trailing version 1.1 without tools and moving ahead once tools are enabled. For teams buying or building agents, that shifts scrutiny toward the whole model-runtime configuration, including its containment controls, rather than the model checkpoint alone.

Meta says Muse Spark 1.2 gains 12.2 points when given tools

Alexandr Wang's Meta Superintelligence Labs, Meta's internal AI organization, reported on August 20th that Muse Spark 1.2 scored 72.0 across 10 multimodal evaluations when it could use tools, up from 59.8 without them. The newly published comparison puts a number on MSL's effort to train the model alongside the runtime in which it works. (research.meta.ai)

AI at Meta on X

That approach fits Wang's history. He co-founded Scale AI at 19 around the thesis that better data and evaluation infrastructure would determine how quickly AI systems improved. Meta recruited him in 2025 after making a $14.3 billion investment in Scale, taking a roughly 49% stake in a transaction that valued Scale at about $29 billion. Wang now serves as Meta's chief AI officer and leads MSL, the organization that rebuilt Meta's model stack and published the first Muse Spark in April. (about.fb.com)

Muse Spark 1.2 arrived on August 5th as the model behind Muse Code, Meta's terminal coding agent. The August 20th technical post reports how the model performs when it can inspect visual material and carry those observations into later reasoning or tool calls. That evaluation is the new element here; Meta had already introduced the underlying model and coding agent earlier in the month.

The improvement appears when the tools arrive

Meta's comparison contains an unusually revealing detail. Muse Spark 1.2 scored 59.8 without tools, slightly below Muse Spark 1.1's 60.2. With tools enabled, 1.2 rose to 72.0, while 1.1 reached 69.1.

That gives version 1.2 a 12.2-point tool-enabled increase, compared with an 8.9-point increase for its predecessor. The newer model leads 1.1 by 2.9 points with tools and trails it by 0.4 points without them. Meta's numbers locate the reported improvement in the interaction between the model and its working environment. (research.meta.ai)

The aggregate covers SimpleVQA, WorldVQA, CharXiv Reasoning, ChartMuseum, ERQA, ChartQA-Pro, OmniSpatial, ZeroBench, BabyVision and PerceptionBench. Meta's public chart provides the average scores, but no per-benchmark results. It also does not establish that an outside evaluator can reproduce the gain. The 12.2-point increase had not been independently replicated as of August 20th, according to Artificial Analysis.

Meta has published a multimodal evaluation methodology alongside the results. The headline figure remains a Meta-run composite, so the useful comparison is the narrow one: two generations of Meta's model measured under Meta's framework, with and without tool access.

Wang's lab is co-training the model and its workplace

Muse Spark 1.2 was co-trained with Muse Code, giving the model experience with the harness it encounters during use. Meta says the training incorporated sampled harness trajectories, goals, context compaction, subagents and the Muse Code toolset.

Muse Code keeps specialist background agents active throughout a session instead of creating a new one for each task. It also appends model calls, tool runs, approvals and edits to a local event log, allowing interrupted work to resume from the recorded state. Those runtime choices target a persistent weakness in coding agents: a capable model can still gather the same context twice, lose an earlier decision or fail during a long task because its surrounding software cannot recover cleanly.

The tool-use score gives Meta a measurable result for that model-and-runtime pairing. It does not show how much of the gain comes from the model, the available tools, the orchestration layer or their combined configuration. Meta has not disclosed per-benchmark scores that would show which visual reasoning skills account for the increase.

Tool use raises the containment stakes

The evaluation arrived six days after Meta disclosed that a misconfigured third-party testing environment allowed a pre-release version of Muse Spark 1.1 to reach the open internet. The model was given the name of a real website instead of a fictional target, found a vulnerability, accessed information and changed the site's database.

Meta said the incident was isolated, involved the earlier 1.1 model and resulted from the evaluator's setup rather than a sandbox escape. The episode supplies a concrete example of the operational risk behind tool-use gains. Longer autonomous workflows and better tool selection demand tighter control over an agent's permissions, targets and network access. (research.meta.ai)

Meta says Muse Spark 1.2 is available through Muse Code and the Meta Model API. The multimodal post also describes the results as arriving ahead of an open-weights release, without setting a date or license. Meta separately released Muse Glimmer, a 30-billion-parameter model intended to run locally, on August 10th.

For Wang, the 1.2 evaluation turns a familiar infrastructure thesis into a model strategy. The direct-response score barely changed across generations. Meta reported a larger gain after connecting Muse Spark 1.2 to the tools and runtime it was trained to use.

Reader comments

Conversation for this story loads after sign-in.