Thứ Năm, 3 thg 9, 2026

Mô hình coding và agentic mới hiện đã có mặt trên Muse Code và Meta Model API. Meta khẳng định đây là bước nhảy vọt lớn nhất của hãng trong mảng lập trình và tác vụ agent. Theo Artificial Analysis, phiên bản xhigh đạt 64 điểm trên Coding Agent Index với chi phí chỉ 1,72 USD mỗi tác vụ, mức giá rẻ nhất đối với bất kỳ agent nào đạt điểm trên 60. Một phiên bản max giới hạn dành cho đối tác của Meta đạt 68 điểm, đứng thứ hai sau Claude Opus 5. So với phiên bản 1.2, Muse Spark 1.3 dùng ít hơn khoảng 20% lượt gọi tool và giảm 25% lượng token trong các thử nghiệm nội bộ. Các benchmark khác cho thấy DeepSWE đạt 75,4 điểm so với 55,0 của Muse Spark 1.2, trong khi Terminal-bench 2.1 đạt 88,8 điểm, ngang ngửa GPT-5.6 Sol. Meta cho biết sẽ sớm ra mắt phiên bản max reasoning và bản open weights của Muse Spark.

Đăng nhập để góp ý, chỉnh sửa
Nguồn chính
SOURCE@shishirpatil_0:27 3 thg 9Muse Spark 1.3 is rolling out today with frontier performance almost too cheap to meter.
SUPPORT@arthurmacwaters0:59 3 thg 9impressive performance from muse spark 1.3, especially given it's a fraction of the cost of fable
SUPPORT@artificialanlys1:07 3 thg 9Meta's Muse Spark 1.3 (max), which is in limited preview for Meta's partners, scores 68 on the Artificial Analysis Coding Agent Index in the Muse Code harness, #2 behind only Claude Opus 5 (xhigh) in Claude Code.
SUPPORT@artificialanlys1:07 3 thg 9it averages ~14M total tokens per task, around a third fewer than Claude Opus 5 (xhigh) in Claude Code (~22M) at the same score.
SUPPORT@wesroth4:30 3 thg 9It even ties GPT-5.6 Sol at 88.8 on Terminal-bench 2.1.
SUPPORT@alexandr_wang0:30 3 thg 9Less than $1, Yes, under $1… in Ultra mode
SOURCE@finkd19:26 2 thg 9frontier performance almost too cheap to meter
SUPPORT@alexandr_wang19:30 2 thg 9holds onto requirements well during long-horizon tasks