Meta’s newest multimodal model, Muse Spark 1.2, can convert visual layouts into working code and orchestrate robotic movements in unstructured environments. The model provides audio-visual understanding for enterprise workflows and incorporates visual inputs directly into its reasoning to improve tool use. In the Agent Arena, the model's net improvement score rose to 2.1% from 0.9% in the previous version, with its Bash Recovery performance climbing from 5.9% to 11.4%.

Meta also previewed WildArtifactBench, an internal framework for assessing agents on real-world tasks across various deliverable formats. The company is releasing 10 tasks from the benchmark, which utilizes win rates and Elo scores from preference judges rather than strict rubrics to measure the practical utility of multimodal agents.

Sign in to suggest edits

Key sources

  1. SOURCE@aiatmeta“turning visuals into working code to translating perception into physical action”x.com
  2. SOURCE@aiatmeta“using win rates and Elo scores from human and agentic preference judges rather than strict ground-truth rubrics”x.com
  3. SUPPORT@aiatmeta“The model inspects visual inputs more closely and incorporates what it finds into its reasoning”x.com
  4. SUPPORT@aiatmeta“evaluating correctness based on actual rendering and behavior while using a continuous self-improvement loop”x.com
  5. SUPPORT@aiatmeta“plan sub-tasks for a bimanual robot tidying a desk, distinguishing a hair brush from a makeup brush”x.com
  6. SUPPORT@arena“Muse Spark 1.2 improves on Muse Spark 1.1, more than doubling the Net Improvement score from 0.9% to 2.1%”x.com
  7. SUPPORT@designarena“lands on the price-preference Pareto frontier for all three categories”x.com
Markdown