---
format: "aidr-story-markdown/v1"
id: "383cec3b0955e288d68096beb8e13a3622d89af8dc3af4b638e6c4f793f13a28"
canonical_url: "https://aidr.today/383cec3b?lang=en"
title: "Ox Alpha Scores 58.4% on DeepSWE, Matching Claude Opus 4.8 Performance"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-22T18:50:03.000Z"
category: "Research"
topics: ["benchmark","coding","agent","llm","reasoning"]
source_urls: ["https://huggingnews.com/ai/ox-alpha-scores-584percent-on-deepswe-matching-claude-opus-48-performanc-4992e222","https://x.com/henryzhangumich/status/2091066210721141009","https://x.com/peterwildeford/status/2091195347016044843","https://x.com/TheZachMueller/status/2091175521077854451","https://x.com/scaling01/status/2091193499949432986","https://x.com/apples_jimmy/status/2091079278100386027","https://x.com/teortaxesTex/status/2091075052523393055","https://x.com/himanshustwts/status/2091243094666822099"]
summary: "A comprehensive evaluation of the latest coding model's agentic capabilities revealed a discrepancy between actual performance and initial market rumors. Testing on the full 113 task DeepSWE benchmark showed Ox Alpha scored 58.4%, a result that undercuts the 80% pass rate previously reported by some users. The result places the model's software engineering performance nearly identical to the 59% achieved by Claude Opus 4.8. Observers noted that this level of parity indicates the Chinese AI development has aligned with top Western models as the industry shifts focus toward multimodal agentic work."
---

# Ox Alpha Scores 58\.4% on DeepSWE, Matching Claude Opus 4\.8 Performance

> [Open the canonical story](<https://aidr.today/383cec3b?lang=en>)

**Published:** 2026-08-22T18:50:03.000Z
**Category:** Research
**Topics:** benchmark, coding, agent, llm, reasoning

## Summary

A comprehensive evaluation of the latest coding model's agentic capabilities revealed a discrepancy between actual performance and initial market rumors\. Testing on the full 113 task DeepSWE benchmark showed Ox Alpha scored 58\.4%, a result that undercuts the 80% pass rate previously reported by some users\. The result places the model's software engineering performance nearly identical to the 59% achieved by Claude Opus 4\.8\. Observers noted that this level of parity indicates the Chinese AI development has aligned with top Western models as the industry shifts focus toward multimodal agentic work\.

## Sources

- [Story source](<https://huggingnews.com/ai/ox-alpha-scores-584percent-on-deepswe-matching-claude-opus-48-performanc-4992e222>)
- [Story source](<https://x.com/henryzhangumich/status/2091066210721141009>)
- [Supporting source](<https://x.com/peterwildeford/status/2091195347016044843>)
- [Supporting source](<https://x.com/TheZachMueller/status/2091175521077854451>)
- [Supporting source](<https://x.com/scaling01/status/2091193499949432986>)
- [Story source](<https://x.com/apples_jimmy/status/2091079278100386027>)
- [Supporting source](<https://x.com/teortaxesTex/status/2091075052523393055>)
- [Story source](<https://x.com/himanshustwts/status/2091243094666822099>)

