Nous Research and other API providers are offering free limited-time access to Ox Alpha, a multimodal reasoning model designed for software engineering and long-horizon production workloads. The model features a 1 million token context window and accepts text, image, and video inputs, with a maximum output of approximately 131,000 tokens per response. Nous Portal reports a daily capacity of 1 quadrillion tokens, though other observers have questioned the substantiation of these throughput claims.

Early tests show Ox Alpha generating complex Three.js 3D scenes in single responses exceeding 64,000 tokens without external assets. On a DeepSWE subset, the model recorded an 80% success rate, surpassing 65% for Fable and 52% for Sol. Market observers speculate the stealth release is the GLM 5.3 Flash model from Chinese laboratory Zhipu, citing an identical tokenizer and similar responses to geopolitical queries.

Sign in to suggest edits

Key sources

  1. SOURCE@teknium“capacity for 1 quadrillion tokens per day”x.com
  2. SUPPORT@rohanpaul_ai“One prompt, 64,745 output tokens, and no external assets: Ox Alpha generated the code for an entire Three.js dreamcore 3D scene/world in a single shot”x.com
  3. SUPPORT@teortaxestex“The 100T/day promise (unsubstantiated so far afaik) is lowkey genius”x.com
  4. SUPPORT@teortaxestex“it does not take a lot of intelligence to perform SWE work at a superhuman level”x.com
  5. SUPPORT@alexatallah“$ ori [your fav harness] --model stealth/ox-alpha”x.com
  6. SUPPORT@askvenice“This is an experimental release. Ox Alpha is free for all users.”x.com
  7. SUPPORT@cline“Early benchmarks shows marginal improvement over Fable and GPT.”x.com
  8. SUPPORT@nousresearch“We have capacity for 1 quadrillion tokens per day.”x.com
Markdown