Meta released Muse Spark 1.2 on August 5, 2026, a coding-focused update shipped alongside Muse Code, Meta's first terminal coding agent. On independent benchmarks, Muse Spark 1.2 scores 54 on Artificial Analysis' Intelligence Index at xhigh reasoning, placing it tied with Grok 4.5 and narrowly behind Opus 5 (61), Fable 5 (60), and GPT-5.6 Sol (59)—a 3-point gain from Muse Spark 1.1 (51) and 11 points from the original Muse Spark (43 in April).
On Meta's own internal benchmarks, the model reached 82.9% on Terminal-Bench 2.1 (up from 1.1's 76.2%) and 59.3% on DeepSWE v1.1. Meta also claimed gold-medal-level performance in five STEM Olympiads: perfect scores on the Asian Physics Olympiad (APhO) and International Physics Olympiad (IPhO) theory exams, plus gold medals on the International Mathematical Olympiad (IMO), International Chemistry Olympiad (IChO), and Romanian Masters of Mathematics (RMM). Three competitions used official live judging with no tools allowed—only pure reasoning through multi-agent orchestration.
On Vals' pricing index, Muse Spark 1.2 entered the top 5 at $0.69/test, reportedly 3x cheaper than Kimi and 10x+ cheaper than Fable, Opus, and GPT-5.6 Sol. Vals also reported it became the first model above 60% on Finance Agent v2 at $0.77/test (versus Opus 5 at $5.12/test) while running 2x the speed. Muse Spark 1.2 maintains standard API pricing of $1.25 per million input tokens and $4.25 per million output tokens, unchanged from 1.1. Meta also offers a contributor tier at 12.5x cheaper input and 21.25x cheaper output in exchange for training-data consent.
For architects, Muse Spark 1.2 signals Meta's four-week release cadence is outpacing the evaluation ecosystem—there is still no published per-benchmark breakdown for 1.2. The Olympiad results hinge on multi-agent orchestration and rejection sampling, not raw model scale, suggesting the path to frontier performance increasingly runs through harnesses and TTC (time-to-completion), not just model weights. The price-to-performance delta matters for cost-sensitive agentic workloads, especially if the Finance Agent v2 setup (2x speed at 1/7th the cost) holds in production.