Meta Superintelligence Labs released Muse Spark 1.2 and Muse Code (beta) on August 5, 2026. Muse Spark 1.2 is a reasoning model explicitly co-trained with Muse Code, Meta's terminal coding agent. Meta claims 82.9% on Terminal-Bench 2.1 and 59% on DeepSWE 1.1. Independent evaluators (Artificial Analysis, Vals) benchmark Muse Spark 1.2 at 54 on the Intelligence Index, tied with Grok 4.5 and 6”7 points behind frontier models (Claude Opus 5: 61, Claude Fable 5: 60, GPT-5.6 Sol: 59). Vals reports $0.69/test on Terminal-Bench using Terminus 2 (common harness)—the lowest cost among its top-five models. Pricing: $1.25 input / $4.25 output per million tokens; Contributor tier at $0.10/$0.20 for data-use permission.
Muse Code implements four agent types (Coordinator, Explorer, Executor, Verifier) communicating through a shared event log that survives restarts without re-deriving context. The terminal agent supports 24–30+ hour jobs (vs typical 10–15 min API call limits), persistent async background agents, and Git-worktree isolation. Meta's internal coding benchmark shows Muse Spark 1.2 at 70.6%, behind Claude Opus 5 (79.4%) but ahead of GPT-5.6 Terra (65.4%). Caveat: the 82.9% is from Meta's own harness; independent Terminal-Bench has no verified entry yet; Meta's Muse Spark 1.1 claim of 80% came in at 76.2% verified, a 3.8-point gap.
For practitioners evaluating long-horizon coding: Muse Code's event-log restart safety and parallel subagents are novel for an open model-side offering. Cost per task matters if you run high volume; Vals' $0.69 undercuts GPT-5.6 and Claude. Meta showed confidence by publishing charts where Claude Opus 5 wins—unusual for a vendor launch. Watch whether the co-training premium (model + harness) holds in other frameworks; if not, Muse Spark 1.2 alone is tier-2 on reasoning but competitive on price.