---
format: "aidr-story-markdown/v1"
id: "46e386fc7356c5c7f3986173ddfa70c04ec2edbabdb9951167d788a7411946f4"
canonical_url: "https://aidr.today/46e386fc?lang=en"
title: "OpenAI Hits 300 Tokens per Second Using Nvidia GPUs Instead of Cerebras"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-09-30T02:37:46.000Z"
category: "Infra"
topics: ["openai","inference"]
source_urls: ["https://marketbrief.now/ai/update-openai-hits-300-tokens-per-second-using-nvidia-gpus-instead-of-ce-66468251","https://huggingnews.com/ai/update-openai-hits-300-tokens-per-second-using-nvidia-gpus-instead-of-ce-66468251"]
summary: "The company recently launched a premium speed tier for GPT-6 Astra that accelerates token generation up to 8x. Analysts at SemiAnalysis believe this performance, which reaches 300 tokens per second, is powered by Nvidia Blackwell-class hardware such as GB300 GPUs running at very low batch sizes. This deployment is a reversal from expectations that OpenAI would use Cerebras hardware for the increase, as a promised 750 tokens per second version from Cerebras remains unavailable. The \"Ultrafast\" mode is available to Enterprise clients, API developers, and individuals through a new $500 per month subscription. OpenAI is expanding the feature to its GPT-6.1 Sol model in the coming weeks. To increase utility for third parties, the company introduced a single sign-on option that allows users to apply their subscription token limits to external applications through their own inference."
---

# OpenAI Hits 300 Tokens per Second Using Nvidia GPUs Instead of Cerebras

> [Open the canonical story](<https://aidr.today/46e386fc?lang=en>)

**Published:** 2026-09-30T02:37:46.000Z
**Category:** Infra
**Topics:** openai, inference

## Summary

The company recently launched a premium speed tier for GPT\-6 Astra that accelerates token generation up to 8x\. Analysts at SemiAnalysis believe this performance, which reaches 300 tokens per second, is powered by Nvidia Blackwell\-class hardware such as GB300 GPUs running at very low batch sizes\. This deployment is a reversal from expectations that OpenAI would use Cerebras hardware for the increase, as a promised 750 tokens per second version from Cerebras remains unavailable\. The "Ultrafast" mode is available to Enterprise clients, API developers, and individuals through a new $500 per month subscription\. OpenAI is expanding the feature to its GPT\-6\.1 Sol model in the coming weeks\. To increase utility for third parties, the company introduced a single sign\-on option that allows users to apply their subscription token limits to external applications through their own inference\.

## Sources

- [Story source](<https://marketbrief.now/ai/update-openai-hits-300-tokens-per-second-using-nvidia-gpus-instead-of-ce-66468251>)
- [Story source](<https://huggingnews.com/ai/update-openai-hits-300-tokens-per-second-using-nvidia-gpus-instead-of-ce-66468251>)

