A comprehensive evaluation of the latest coding model's agentic capabilities revealed a discrepancy between actual performance and initial market rumors. Testing on the full 113 task DeepSWE benchmark showed Ox Alpha scored 58.4%, a result that undercuts the 80% pass rate previously reported by some users.

The result places the model's software engineering performance nearly identical to the 59% achieved by Claude Opus 4.8. Observers noted that this level of parity indicates the Chinese AI development has aligned with top Western models as the industry shifts focus toward multimodal agentic work.

Sign in to suggest edits

Key sources

  1. SOURCE@henryzhangumich“Actual benchmark result is 58.4%, landing the model at almost identical performance to Claude Opus 4.8 (59%)”x.com
  2. SUPPORT@peterwildeford“China should have an Opus 4.8 model about now”x.com
  3. SUPPORT@thezachmueller“It’s Opus 4.8, disappointing!”x.com
  4. SUPPORT@scaling01“The Rumored ~80% pass rate is completely incorrect.”x.com
  5. SOURCE@apples_jimmy“Just completed a full end to end run of the full 113 task DeepSWE benchmark on ox-alpha.”x.com
  6. SUPPORT@teortaxestex“The 100T/day promise (unsubstantiated so far afaik) is lowkey genius”x.com
  7. SOURCE@himanshustwts“Text input follows the same format as glm-5.3 and multimodal input follows the same format as glm-5v-turbo”x.com
  8. SUPPORT@andrewcurran_“possibly a smaller variant of the new flagship large model they are still training”x.com
Markdown