Mô hình coding và agentic mới hiện đã có mặt trên Muse Code và Meta Model API. Meta khẳng định đây là bước nhảy vọt lớn nhất của hãng trong mảng lập trình và tác vụ agent. Theo Artificial Analysis, phiên bản xhigh đạt 64 điểm trên Coding Agent Index với chi phí chỉ 1,72 USD mỗi tác vụ, mức giá rẻ nhất đối với bất kỳ agent nào đạt điểm trên 60. Một phiên bản max giới hạn dành cho đối tác của Meta đạt 68 điểm, đứng thứ hai sau Claude Opus 5. So với phiên bản 1.2, Muse Spark 1.3 dùng ít hơn khoảng 20% lượt gọi tool và giảm 25% lượng token trong các thử nghiệm nội bộ. Các benchmark khác cho thấy DeepSWE đạt 75,4 điểm so với 55,0 của Muse Spark 1.2, trong khi Terminal-bench 2.1 đạt 88,8 điểm, ngang ngửa GPT-5.6 Sol. Meta cho biết sẽ sớm ra mắt phiên bản max reasoning và bản open weights của Muse Spark.

Đăng nhập để góp ý, chỉnh sửa

Nguồn chính

  1. SOURCE@shishirpatil_“Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter.”x.com
  2. SUPPORT@arthurmacwaters“impressive performance from muse spark 1.3, especially given it's a fraction of the cost of fable”x.com
  3. SUPPORT@artificialanlys“Meta's Muse Spark 1.3 (max), which is in limited preview for Meta's partners, scores 68 on the Artificial Analysis Coding Agent Index in the Muse Code harness, #2 behind only Claude Opus 5 (xhigh) in Claude Code.”x.com
  4. SUPPORT@artificialanlys“it averages ~14M total tokens per task, around a third fewer than Claude Opus 5 (xhigh) in Claude Code (~22M) at the same score.”x.com
  5. SUPPORT@wesroth“It even ties GPT-5.6 Sol at 88.8 on Terminal-bench 2.1.”x.com
  6. SUPPORT@alexandr_wang“Less than $1, Yes, under $1… in Ultra mode”x.com
  7. SOURCE@finkd“frontier performance almost too cheap to meter”x.com
  8. SUPPORT@alexandr_wang“holds onto requirements well during long-horizon tasks”x.com
Bản Markdown