The new coding and agentic model is available now in Muse Code and the Meta Model API, and Meta said it is its biggest jump yet in coding and agentic work. Artificial Analysis said the xhigh version scored 64 on its Coding Agent Index at $1.72 per task, the lowest cost of any agent above a 60 score. A limited-preview max version for Meta's partners scored 68 and ranked second behind Claude Opus 5.

Meta said Muse Spark 1.3 uses about 20% fewer tool calls and about 25% fewer tokens than Muse Spark 1.2 in internal comparisons. Separate benchmark results put DeepSWE at 75.4 versus 55.0 for Muse Spark 1.2 and Terminal-bench 2.1 at 88.8, matching GPT-5.6 Sol. Meta said max reasoning and a Muse Spark open weights release are coming later.

Sign in to suggest edits

Key sources

  1. SOURCE@shishirpatil_“Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter.”x.com
  2. SUPPORT@arthurmacwaters“impressive performance from muse spark 1.3, especially given it's a fraction of the cost of fable”x.com
  3. SUPPORT@artificialanlys“Meta's Muse Spark 1.3 (max), which is in limited preview for Meta's partners, scores 68 on the Artificial Analysis Coding Agent Index in the Muse Code harness, #2 behind only Claude Opus 5 (xhigh) in Claude Code.”x.com
  4. SUPPORT@artificialanlys“it averages ~14M total tokens per task, around a third fewer than Claude Opus 5 (xhigh) in Claude Code (~22M) at the same score.”x.com
  5. SUPPORT@wesroth“It even ties GPT-5.6 Sol at 88.8 on Terminal-bench 2.1.”x.com
  6. SUPPORT@alexandr_wang“Less than $1, Yes, under $1… in Ultra mode”x.com
  7. SOURCE@finkd“frontier performance almost too cheap to meter”x.com
  8. SUPPORT@alexandr_wang“holds onto requirements well during long-horizon tasks”x.com
Markdown