Thứ Bảy, 22 thg 8, 2026

Đánh giá toàn diện về khả năng agentic của mô hình lập trình mới nhất cho thấy có sự khác biệt giữa hiệu suất thực tế và các tin đồn trước đó trên thị trường. Kết quả kiểm tra trên toàn bộ 113 tác vụ của benchmark DeepSWE cho thấy Ox Alpha đạt 58,4%, thấp hơn mức tỷ lệ vượt qua 80% mà một số người dùng từng báo cáo.

Kết quả này đưa hiệu suất kỹ thuật phần mềm của mô hình đạt mức gần như tương đương với con số 59% của Claude Opus 4.8. Các chuyên gia nhận định sự tương đồng này cho thấy AI Trung Quốc đang dần bắt kịp các mô hình hàng đầu phương Tây khi ngành công nghiệp chuyển trọng tâm sang các agent đa phương thức.

Đăng nhập để góp ý, chỉnh sửa
Nguồn chính
SOURCE@henryzhangumich7:33 22 thg 8Actual benchmark result is 58.4%, landing the model at almost identical performance to Claude Opus 4.8 (59%)
SUPPORT@peterwildeford16:06 22 thg 8China should have an Opus 4.8 model about now
SUPPORT@thezachmueller14:47 22 thg 8It’s Opus 4.8, disappointing!
SUPPORT@scaling0115:59 22 thg 8The Rumored ~80% pass rate is completely incorrect.
SOURCE@apples_jimmy8:25 22 thg 8Just completed a full end to end run of the full 113 task DeepSWE benchmark on ox-alpha.
SUPPORT@teortaxestex8:08 22 thg 8The 100T/day promise (unsubstantiated so far afaik) is lowkey genius
SOURCE@himanshustwts19:16 22 thg 8Text input follows the same format as glm-5.3 and multimodal input follows the same format as glm-5v-turbo
SUPPORT@andrewcurran_20:13 22 thg 8possibly a smaller variant of the new flagship large model they are still training