---
format: "aidr-story-markdown/v1"
id: "beb40dada37bee37f131850f3b9cb2b175e3798d2c3a3fd06721aadff6079d86"
canonical_url: "https://aidr.today/beb40dad?lang=en"
title: "Ox Alpha Model Hits 1 Quadrillion Daily Token Capacity to Beat Fable Coding SOTA"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-22T08:48:35.000Z"
category: "Models"
topics: ["coding","reasoning","multimodal","benchmark"]
source_urls: ["https://x.com/Teknium/status/2090979852879073629","https://x.com/rohanpaul_ai/status/2090967376879870462","https://x.com/teortaxesTex/status/2091075052523393055","https://x.com/teortaxesTex/status/2090966752481972319","https://x.com/alexatallah/status/2090835555151925703","https://x.com/AskVenice/status/2090847078633156947","https://x.com/cline/status/2090854216399220985","https://x.com/NousResearch/status/2090899914700054780"]
summary: "Nous Research and other API providers are offering free limited-time access to Ox Alpha, a multimodal reasoning model designed for software engineering and long-horizon production workloads. The model features a 1 million token context window and accepts text, image, and video inputs, with a maximum output of approximately 131,000 tokens per response. Nous Portal reports a daily capacity of 1 quadrillion tokens, though other observers have questioned the substantiation of these throughput claims. Early tests show Ox Alpha generating complex Three.js 3D scenes in single responses exceeding 64,000 tokens without external assets. On a DeepSWE subset, the model recorded an 80% success rate, surpassing 65% for Fable and 52% for Sol. Market observers speculate the stealth release is the GLM 5.3 Flash model from Chinese laboratory Zhipu, citing an identical tokenizer and similar responses to geopolitical queries."
---

# Ox Alpha Model Hits 1 Quadrillion Daily Token Capacity to Beat Fable Coding SOTA

> [Open the canonical story](<https://aidr.today/beb40dad?lang=en>)

**Published:** 2026-08-22T08:48:35.000Z
**Category:** Models
**Topics:** coding, reasoning, multimodal, benchmark

## Summary

Nous Research and other API providers are offering free limited\-time access to Ox Alpha, a multimodal reasoning model designed for software engineering and long\-horizon production workloads\. The model features a 1 million token context window and accepts text, image, and video inputs, with a maximum output of approximately 131,000 tokens per response\. Nous Portal reports a daily capacity of 1 quadrillion tokens, though other observers have questioned the substantiation of these throughput claims\. Early tests show Ox Alpha generating complex Three\.js 3D scenes in single responses exceeding 64,000 tokens without external assets\. On a DeepSWE subset, the model recorded an 80% success rate, surpassing 65% for Fable and 52% for Sol\. Market observers speculate the stealth release is the GLM 5\.3 Flash model from Chinese laboratory Zhipu, citing an identical tokenizer and similar responses to geopolitical queries\.

## Sources

- [Story source](<https://x.com/Teknium/status/2090979852879073629>)
- [Supporting source](<https://x.com/rohanpaul_ai/status/2090967376879870462>)
- [Supporting source](<https://x.com/teortaxesTex/status/2091075052523393055>)
- [Supporting source](<https://x.com/teortaxesTex/status/2090966752481972319>)
- [Supporting source](<https://x.com/alexatallah/status/2090835555151925703>)
- [Supporting source](<https://x.com/AskVenice/status/2090847078633156947>)
- [Supporting source](<https://x.com/cline/status/2090854216399220985>)
- [Supporting source](<https://x.com/NousResearch/status/2090899914700054780>)

