Coding agents leveraging repeated tool calls and long contexts drove a surge in token consumption for the newly released Ox Alpha model. The model processed 26 trillion tokens in four days on the OpenRouter platform, a volume 2.6 times larger than any previous model launch on the service. To support sustained agentic work, Ox Alpha utilizes a 1.05M token context window that allows sessions to carry more text than typical chat applications.

Single-day usage for the model reached 5 trillion tokens, roughly 30 times the daily volume of Claude Opus 5 or GPT-5.6 Sol. The model is currently offered for free via the ori harness to encourage adoption among developers using coding agents.

Sign in to suggest edits

Key sources

  1. SUPPORT@rohanpaul_ai“coding agents, where long contexts, retries, and repeated tool calls can massively multiply token consumption inside one task”x.com
  2. SUPPORT@ycombinator“26T tokens of Ox Alpha in 4 days”x.com
  3. SUPPORT@nicbstme“Intelligence really is becoming too cheap to meter (the model is free right now😂)”x.com
  4. SOURCEhuggingnewshuggingnews.com
  5. SOURCE@vipulved“0x Alpha will be on @togethercompute as early as tomorrow morning!”x.com
  6. SOURCE@aaazzam“Ox Alpha, which is coming to Modal tomorrow”x.com
  7. SUPPORT@togethercompute“Excited to share that 0x Alpha will be on @togethercompute”x.com
  8. SOURCEhuggingnewshuggingnews.com
Markdown