---
format: "aidr-story-markdown/v1"
id: "e50cc370a44f0882cd10fbc567eb5de84f04e0025a2c891e320d268b89d39654"
canonical_url: "https://aidr.today/e50cc370?lang=en"
title: "Meta Muse Spark 1.2 More Than Doubles Agent Arena Improvement to 2.1%"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-20T23:55:06.000Z"
category: "Models"
topics: ["meta","agent","multimodal","coding","benchmark"]
source_urls: ["https://huggingnews.com/ai/meta-muse-spark-12-more-than-doubles-agent-arena-improvement-to-21percen-9305d7e9","https://x.com/AIatMeta/status/2090485743034716420","https://x.com/AIatMeta/status/2090505413817246050","https://x.com/AIatMeta/status/2090485747497546004","https://x.com/AIatMeta/status/2090485750076989595","https://x.com/AIatMeta/status/2090485751926706501","https://x.com/arena/status/2090484142408618033","https://x.com/DesignArena/status/2090498670685020639"]
summary: "Meta’s newest multimodal model, Muse Spark 1.2, can convert visual layouts into working code and orchestrate robotic movements in unstructured environments. The model provides audio-visual understanding for enterprise workflows and incorporates visual inputs directly into its reasoning to improve tool use. In the Agent Arena, the model's net improvement score rose to 2.1% from 0.9% in the previous version, with its Bash Recovery performance climbing from 5.9% to 11.4%. Meta also previewed WildArtifactBench, an internal framework for assessing agents on real-world tasks across various deliverable formats. The company is releasing 10 tasks from the benchmark, which utilizes win rates and Elo scores from preference judges rather than strict rubrics to measure the practical utility of multimodal agents."
---

# Meta Muse Spark 1\.2 More Than Doubles Agent Arena Improvement to 2\.1%

> [Open the canonical story](<https://aidr.today/e50cc370?lang=en>)

**Published:** 2026-08-20T23:55:06.000Z
**Category:** Models
**Topics:** meta, agent, multimodal, coding, benchmark

## Summary

Meta’s newest multimodal model, Muse Spark 1\.2, can convert visual layouts into working code and orchestrate robotic movements in unstructured environments\. The model provides audio\-visual understanding for enterprise workflows and incorporates visual inputs directly into its reasoning to improve tool use\. In the Agent Arena, the model's net improvement score rose to 2\.1% from 0\.9% in the previous version, with its Bash Recovery performance climbing from 5\.9% to 11\.4%\. Meta also previewed WildArtifactBench, an internal framework for assessing agents on real\-world tasks across various deliverable formats\. The company is releasing 10 tasks from the benchmark, which utilizes win rates and Elo scores from preference judges rather than strict rubrics to measure the practical utility of multimodal agents\.

## Sources

- [Story source](<https://huggingnews.com/ai/meta-muse-spark-12-more-than-doubles-agent-arena-improvement-to-21percen-9305d7e9>)
- [Story source](<https://x.com/AIatMeta/status/2090485743034716420>)
- [Story source](<https://x.com/AIatMeta/status/2090505413817246050>)
- [Supporting source](<https://x.com/AIatMeta/status/2090485747497546004>)
- [Supporting source](<https://x.com/AIatMeta/status/2090485750076989595>)
- [Supporting source](<https://x.com/AIatMeta/status/2090485751926706501>)
- [Supporting source](<https://x.com/arena/status/2090484142408618033>)
- [Supporting source](<https://x.com/DesignArena/status/2090498670685020639>)

