The company recently launched a premium speed tier for GPT-6 Astra that accelerates token generation up to 8x. Analysts at SemiAnalysis believe this performance, which reaches 300 tokens per second, is powered by Nvidia Blackwell-class hardware such as GB300 GPUs running at very low batch sizes. This deployment is a reversal from expectations that OpenAI would use Cerebras hardware for the increase, as a promised 750 tokens per second version from Cerebras remains unavailable.

The "Ultrafast" mode is available to Enterprise clients, API developers, and individuals through a new $500 per month subscription. OpenAI is expanding the feature to its GPT-6.1 Sol model in the coming weeks. To increase utility for third parties, the company introduced a single sign-on option that allows users to apply their subscription token limits to external applications through their own inference.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCEhuggingnewshuggingnews.com
Markdown