---
format: "aidr-story-markdown/v1"
id: "8828f586f78075c7a1f1df4029b59215fc1472eb1b04ee0df6f5b710c679babb"
canonical_url: "https://aidr.today/8828f586?lang=en"
title: "DeepSeek v4 Flash Token Volume Halves, Ending 18T Daily Peak After 5x Price Hike"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-21T08:46:12.000Z"
category: "Industry"
topics: ["deepseek","inference","llm","vision","open-source"]
source_urls: ["https://x.com/jayair/status/2090596382306361380","https://x.com/teortaxesTex/status/2090702916847849812","https://news.ycombinator.com/item?id=49386163","https://api-docs.deepseek.com/guides/vision/"]
summary: "DeepSeek v4 Flash token usage on OpenCode fell sharply after the AI company increased retail pricing for the model's lightweight version. The firm raised rates fivefold for peak hours and 2.5x for off-peak hours on August 16, sending daily volumes to less than half of a recent peak. Usage had previously surged from 3T to 18T tokens a day within two weeks of the model's August 1 launch. The explosive growth was driven by pricing that was 350x cheaper than Fable and 70x cheaper than Sonnet, supported by custom infrastructure to cache tokens. The price hike follows a 6x jump in demand that strained capacity and coincided with rising GPU costs. This volume drop has created a nearly 10T token per day void in the market for low-cost inference, which competing model labs may seek to fill."
---

# DeepSeek v4 Flash Token Volume Halves, Ending 18T Daily Peak After 5x Price Hike

> [Open the canonical story](<https://aidr.today/8828f586?lang=en>)

**Published:** 2026-08-21T08:46:12.000Z
**Category:** Industry
**Topics:** deepseek, inference, llm, vision, open\-source

## Summary

DeepSeek v4 Flash token usage on OpenCode fell sharply after the AI company increased retail pricing for the model's lightweight version\. The firm raised rates fivefold for peak hours and 2\.5x for off\-peak hours on August 16, sending daily volumes to less than half of a recent peak\. Usage had previously surged from 3T to 18T tokens a day within two weeks of the model's August 1 launch\. The explosive growth was driven by pricing that was 350x cheaper than Fable and 70x cheaper than Sonnet, supported by custom infrastructure to cache tokens\. The price hike follows a 6x jump in demand that strained capacity and coincided with rising GPU costs\. This volume drop has created a nearly 10T token per day void in the market for low\-cost inference, which competing model labs may seek to fill\.

## Sources

- [Story source](<https://x.com/jayair/status/2090596382306361380>)
- [Supporting source](<https://x.com/teortaxesTex/status/2090702916847849812>)
- [Discussion](<https://news.ycombinator.com/item?id=49386163>)
- [Story source](<https://api-docs.deepseek.com/guides/vision/>)

