---
format: "aidr-story-markdown/v1"
id: "a48e0d6684c43d9c88b3f32ff6298d7f12d9435ad66b59a5e487ea78b1f4bcdf"
canonical_url: "https://aidr.today/a48e0d66?lang=en"
title: "Ox Alpha Hits 80% on Full DeepSWE Benchmark, First Run of All 113 Tasks"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-22T14:48:07.000Z"
category: "Models"
topics: ["llm","coding","benchmark","agent"]
source_urls: ["https://huggingnews.com/ai/update-ox-alpha-hits-80percent-on-full-deepswe-benchmark-first-run-of-al-18a12436","https://x.com/apples_jimmy/status/2091079278100386027","https://x.com/Teknium/status/2090979852879073629","https://x.com/rohanpaul_ai/status/2090967376879870462","https://x.com/teortaxesTex/status/2091075052523393055","https://x.com/teortaxesTex/status/2090966752481972319"]
summary: "A comprehensive evaluation of the DeepSWE coding benchmark confirms that the stealth AI model Ox Alpha performs at the high levels suggested by previous limited tests. The model finished a complete run of 113 tasks with a pass rate of approximately 80%, which researcher Henry Zhang reported as a validation of the model's agentic capabilities. Ox Alpha is currently available for free on the OpenRouter, Venice, and Nous Research platforms to gather feedback on its production utility. The model includes a 1M token context window and supports multimodal input across text, image, and video. Early users suspect the model is GLM 5.3 Flash from the Chinese lab Zhipu AI because of identical tokenization and nearly identical responses to political queries. In earlier subset tests, Ox Alpha's 80% score surpassed results of 65% for Fable and 52% for Sol."
---

# Ox Alpha Hits 80% on Full DeepSWE Benchmark, First Run of All 113 Tasks

> [Open the canonical story](<https://aidr.today/a48e0d66?lang=en>)

**Published:** 2026-08-22T14:48:07.000Z
**Category:** Models
**Topics:** llm, coding, benchmark, agent

## Summary

A comprehensive evaluation of the DeepSWE coding benchmark confirms that the stealth AI model Ox Alpha performs at the high levels suggested by previous limited tests\. The model finished a complete run of 113 tasks with a pass rate of approximately 80%, which researcher Henry Zhang reported as a validation of the model's agentic capabilities\. Ox Alpha is currently available for free on the OpenRouter, Venice, and Nous Research platforms to gather feedback on its production utility\. The model includes a 1M token context window and supports multimodal input across text, image, and video\. Early users suspect the model is GLM 5\.3 Flash from the Chinese lab Zhipu AI because of identical tokenization and nearly identical responses to political queries\. In earlier subset tests, Ox Alpha's 80% score surpassed results of 65% for Fable and 52% for Sol\.

## Sources

- [Story source](<https://huggingnews.com/ai/update-ox-alpha-hits-80percent-on-full-deepswe-benchmark-first-run-of-al-18a12436>)
- [Story source](<https://x.com/apples_jimmy/status/2091079278100386027>)
- [Story source](<https://x.com/Teknium/status/2090979852879073629>)
- [Supporting source](<https://x.com/rohanpaul_ai/status/2090967376879870462>)
- [Supporting source](<https://x.com/teortaxesTex/status/2091075052523393055>)
- [Supporting source](<https://x.com/teortaxesTex/status/2090966752481972319>)

