A comprehensive evaluation of the DeepSWE coding benchmark confirms that the stealth AI model Ox Alpha performs at the high levels suggested by previous limited tests. The model finished a complete run of 113 tasks with a pass rate of approximately 80%, which researcher Henry Zhang reported as a validation of the model's agentic capabilities. Ox Alpha is currently available for free on the OpenRouter, Venice, and Nous Research platforms to gather feedback on its production utility.
The model includes a 1M token context window and supports multimodal input across text, image, and video. Early users suspect the model is GLM 5.3 Flash from the Chinese lab Zhipu AI because of identical tokenization and nearly identical responses to political queries. In earlier subset tests, Ox Alpha's 80% score surpassed results of 65% for Fable and 52% for Sol.
Key sources
- SOURCE@apples_jimmy“Just completed a full end to end run of the full 113 task DeepSWE benchmark on ox-alpha.”x.com
- SOURCE@teknium“capacity for 1 quadrillion tokens per day”x.com
- SUPPORT@rohanpaul_ai“One prompt, 64,745 output tokens, and no external assets: Ox Alpha generated the code for an entire Three.js dreamcore 3D scene/world in a single shot”x.com
- SUPPORT@teortaxestex“The 100T/day promise (unsubstantiated so far afaik) is lowkey genius”x.com
- SUPPORT@teortaxestex“it does not take a lot of intelligence to perform SWE work at a superhuman level”x.com