Meta's Muse Spark 1.3 gets faster when pushed harder, still passes just 14 of 23 tasks

Morgan Linton's 92-run sweep found lower effort made Meta's model slower and less accurate, while Astra and Fable 5.1 fully passed all 23 tasks.

By · Published

Primary source: Morgan Linton on X

Why it matters

VulcanBench's results show why teams cannot route coding work from max-effort launch scores alone. Muse Spark 1.3 stayed cheap, but its lower settings cost accuracy and time.

A close-up view of a high-performance server rack in a modern data lab, with abstract data patterns displayed on a monitor in the foreground.

Morgan Linton (@morganlinton), the creator of VulcanBench, found that Meta's Muse Spark 1.3 fully passed 14 of 23 coding tasks at its best tested effort level, trailing GPT-6 Astra and Claude Fable 5.1 on the same hidden-test suite.

Linton published the results in a five-post thread on September 21st, calling it the longest-running benchmark he has completed for VulcanBench. The 92-run sweep covered four effort settings in Meta's Muse Code harness: low, medium, high and extra-high. The contributor plan used for the test did not offer a max setting.

Muse Spark 1.3's best functional score was 76.87 at extra-high effort, where it fully passed 14 tasks. Astra and Fable 5.1 each fully passed all 23 tasks at their best settings, producing functional scores of 100. Astra reached that mark at medium effort in 3.8 minutes per task. Fable reached it at max effort in 27.1 minutes per task.

Muse took an average of 36 minutes per task at extra-high. Its lower settings were slower: about 70 minutes at low, 69 minutes at medium and 54 minutes at high. A minimal-effort batch, omitted from the main chart because of its scale, averaged 134 minutes per task and scored 71.7.

That inverted runtime curve is the central result. Meta's model became faster as VulcanBench increased the reasoning setting, while its functional score also rose. Linton described the behavior as an apparent "overthinking issue" and said Fable or Astra at low effort would be preferable to Muse Spark 1.3 at any tested setting.

Cheap tokens, expensive waiting

The Muse sweep consumed 2.46 billion raw tokens across 92 runs. VulcanBench calculated an API-equivalent cost of $19.66 using Meta's heavily discounted contributor rates, where prompts and completions may be used to improve Meta's products. The cost ranged from $0.14 to $0.27 per task across the four effort settings.

VulcanBench cost and token comparison for Astra, Fable 5.1 and Muse Spark 1.3
Muse Spark 1.3 consumed 2.46 billion raw tokens across the four tested effort levels. Graphic: VulcanBench.

VulcanBench estimated that the same token volume would cost about $583 at Meta's standard API rates. The comparison does not represent Linton's subscription bill, and it excludes judging costs. It does show the bargain Meta is offering developers willing to contribute their coding data: the discount absorbs a token footprint that would otherwise overwhelm Muse Spark 1.3's apparent price advantage.

The benchmark also carries an important limitation. Astra ran through Codex, Fable through Claude Code and Muse Spark 1.3 through Muse Code. VulcanBench therefore measured each model together with its coding harness, rather than isolating the underlying models behind a common agent loop. Linton used the same 23 tasks and hidden tests for all three, one task at a time, but the harnesses can influence tool use, context management and recovery from errors.

Muse Spark 1.3 also had not received VulcanBench's separate code-quality review when Linton published the chart. Its 76.87 figure is a functional score based on hidden tests, with partial credit available. Astra and Fable have combined scores that include model-judged readability and maintainability, but those combined figures cannot yet be compared directly with Muse.

A benchmark built from an engineering bill

Linton is also co-founder and CTO of Bold Metrics, which builds machine-learning systems for predicting body measurements and apparel sizing. He previously spent nine years at Sonos and holds undergraduate and master's degrees in computer engineering from Carnegie Mellon University.

He started VulcanBench after Bold Metrics' use of coding models was heading toward what he described as a $300,000 annual bill. Linton says testing different models and reasoning settings against his team's work helped reduce that projected spend to less than $30,000 a year. The experience shaped VulcanBench around effort sweeps, cost and completion time rather than a single maximum-effort score.

That design makes the Muse result sharper than a conventional model leaderboard. Meta introduced Muse Spark 1.3 on September 2nd with company-reported gains in coding, agentic work and token efficiency. Independent measurements have supported some of those claims on established coding tests, including Terminal-Bench and long-horizon software engineering evaluations.

Linton's test examines a different purchasing decision: which model and reasoning setting an engineering team should route an actual task to. Muse Spark 1.3 remained inexpensive under Meta's contributor pricing, but the benchmark found no setting where it matched the full-task completion rate of Astra or Fable 5.1. At lower effort levels, developers also gave up time rather than saving it.

Reader comments

Conversation for this story loads after sign-in.