Model facet · Benchmarks

Muse Spark 1.3 Benchmarks: All 11 Launch Scores

Meta reports 11 launch scores for Muse Spark 1.3. The scorecard covers seven agentic tests, three coding tests and one knowledge test, with results reported per benchmark rather than against another model.

Muse Spark 1.3 scorecard

Section headers group the eleven rows by capability. The page intentionally compares Muse Spark 1.3 against itself only; every other model on the site compares against Muse Spark 1.3 instead.

BenchmarkMuse Spark 1.3
Agentic
JobBench64.9
OSWorld 2.066.9
DeepSearchQA89.4
Agentic IF Index57.8
AutomationBench49.4
MRCR 256K–512K98.5
MRCR 512K–1M98.1
Coding
DeepSWE v1.175.4
SWEAtlas CodeBase QnA59.4
Terminal-Bench 2.188.8
Knowledge
GDPval-AA v21754

How to read the scorecard

  • Muse Spark 1.3 leads both MRCR long-context tests (98.5 at 256K–512K, 98.1 at 512K–1M) and the GDPval-AA v2 knowledge score (1754 Elo).
  • Coding results are reported per test: DeepSWE v1.1 (75.4), SWEAtlas CodeBase QnA (59.4) and Terminal-Bench 2.1 (88.8).
  • The Agentic IF Index at 57.8 is a composite, not a single-task score; treat it separately from the individual agentic rows above.
  • Artificial Analysis separately scores Muse Spark 1.3 max at 62 and xhigh at 61 on its Intelligence Index; that composite is not part of the 11-row Meta scorecard.

Ready to test the workflow?

Create account & add credits

Frequently asked questions

Can I run Muse Spark 1.3 on OneInfer?

No. As of 3 September 2026, Muse Spark 1.3 is proprietary, has no downloadable weights, and resolves to Meta as its single upstream provider. OneInfer publishes this independent reference without claiming deployment availability.

How many launch benchmarks did Meta publish for Muse Spark 1.3?

Eleven: seven agentic, three coding and one knowledge.

What is its Artificial Analysis score?

62 for the gated max variant and 61 for the generally available xhigh variant as observed on 3 September 2026.

Put Muse Spark 1.3 to work

Fund a controlled evaluation, send a reference frame or document, and measure quality and cost on your own workload.