aiexpert
Home / News / Brief
Research · Aug 09, 2026, 02:32 AM · 4 sources

Meta Muse Spark 1.2 wins five STEM Olympiad golds, hits $0.69/test on Vals Index

Meta released Muse Spark 1.2 on August 5, 2026, a coding-focused update shipped alongside Muse Code, Meta's first terminal coding agent. On independent benchmarks, Muse Spark 1.2 scores 54 on Artificial Analysis' Intelligence Index at xhigh reasoning, placing it tied with Grok 4.5 and narrowly behind Opus 5 (61), Fable 5 (60), and GPT-5.6 Sol (59)—a 3-point gain from Muse Spark 1.1 (51) and 11 points from the original Muse Spark (43 in April).

On Meta's own internal benchmarks, the model reached 82.9% on Terminal-Bench 2.1 (up from 1.1's 76.2%) and 59.3% on DeepSWE v1.1. Meta also claimed gold-medal-level performance in five STEM Olympiads: perfect scores on the Asian Physics Olympiad (APhO) and International Physics Olympiad (IPhO) theory exams, plus gold medals on the International Mathematical Olympiad (IMO), International Chemistry Olympiad (IChO), and Romanian Masters of Mathematics (RMM). Three competitions used official live judging with no tools allowed—only pure reasoning through multi-agent orchestration.

On Vals' pricing index, Muse Spark 1.2 entered the top 5 at $0.69/test, reportedly 3x cheaper than Kimi and 10x+ cheaper than Fable, Opus, and GPT-5.6 Sol. Vals also reported it became the first model above 60% on Finance Agent v2 at $0.77/test (versus Opus 5 at $5.12/test) while running 2x the speed. Muse Spark 1.2 maintains standard API pricing of $1.25 per million input tokens and $4.25 per million output tokens, unchanged from 1.1. Meta also offers a contributor tier at 12.5x cheaper input and 21.25x cheaper output in exchange for training-data consent.

For architects, Muse Spark 1.2 signals Meta's four-week release cadence is outpacing the evaluation ecosystem—there is still no published per-benchmark breakdown for 1.2. The Olympiad results hinge on multi-agent orchestration and rejection sampling, not raw model scale, suggesting the path to frontier performance increasingly runs through harnesses and TTC (time-to-completion), not just model weights. The price-to-performance delta matters for cost-sensitive agentic workloads, especially if the Finance Agent v2 setup (2x speed at 1/7th the cost) holds in production.

Sources

Everything this brief rests on
  1. 01 Primary source news.smol.ai
  2. 02 artificialanalysis.ai artificialanalysis.ai “Meta's Muse Spark 1.2 scores 54 on the Artificial Analysis Intelligence Index”
  3. 03 x.com x.com “Meta's internal Muse Spark models achieved gold-medal results across the International Physics Olympiad, Asian Physics Olympiad, International Mathematical Olympiad, International Chemistry Olympiad, and Romanian Masters of Mathematics”
  4. 04 orcarouter.ai orcarouter.ai “Muse Spark 1.2 costs $0.40 per Intelligence Index task at Meta's unchanged $1.25/$4.25 per 1M token pricing, with only Grok 4.5 (high, $0.37) and GPT-5.6 Sol (medium, $0.39) cheaper in its intelligence cluster”