---
format: "aidr-story-markdown/v1"
id: "17e4eab7f33c9fac068f7a0a0eeb752bff6747e182ad5a4ad3260f28cb95c257"
canonical_url: "https://aidr.today/17e4eab7?lang=en"
title: "Nvidia Groq 3 LPX Hits 3,400 Tokens Per Second in Nebius Cloud Debut"
lang: "en"
requested_lang: "en"
available_langs: ["en","vi"]
translation_fallback: null
fallback_fields: []
published_at: "2026-08-24T16:52:09.000Z"
category: "Chips"
topics: ["nvidia","inference","agent","chips"]
source_urls: ["https://huggingnews.com/ai/nvidia-groq-3-lpx-hits-3400-tokens-per-second-in-nebius-cloud-debut-9615e020","https://x.com/nvidia/status/2091911896538677689","https://x.com/wallstengine/status/2091903357380342066","https://x.com/FirstSquawk/status/2091903682724216978","https://x.com/FirstSquawk/status/2091903786520572124","https://x.com/CNBC/status/2091904292915388796","https://x.com/JonathanRoss321/status/2091914710325289435"]
summary: "Nvidia has moved its Groq 3 LPX low-latency inference accelerator into full production to support coding agents and multi-step reasoning workloads. Designed to extend the Vera Rubin NVL72 platform, the architecture splits inference work between Rubin GPUs for large-scale context processing and the LPX for fast token generation. In Artificial Analysis tests running Gemma 4 31B with a 100K token context, the chip reached 3,400 output tokens per second, which Nvidia says is 4x faster than the nearest alternative for latency-sensitive agentic tasks. Nebius will be the first AI cloud to deploy the hardware through its Token Factory, with Groq also among the earliest adopters. The rollout follows a $20 billion purchase, and Nvidia says Groq racks will be online by the end of 2026."
---

# Nvidia Groq 3 LPX Hits 3,400 Tokens Per Second in Nebius Cloud Debut

> [Open the canonical story](<https://aidr.today/17e4eab7?lang=en>)

**Published:** 2026-08-24T16:52:09.000Z
**Category:** Chips
**Topics:** nvidia, inference, agent, chips

## Summary

Nvidia has moved its Groq 3 LPX low\-latency inference accelerator into full production to support coding agents and multi\-step reasoning workloads\. Designed to extend the Vera Rubin NVL72 platform, the architecture splits inference work between Rubin GPUs for large\-scale context processing and the LPX for fast token generation\. In Artificial Analysis tests running Gemma 4 31B with a 100K token context, the chip reached 3,400 output tokens per second, which Nvidia says is 4x faster than the nearest alternative for latency\-sensitive agentic tasks\. Nebius will be the first AI cloud to deploy the hardware through its Token Factory, with Groq also among the earliest adopters\. The rollout follows a $20 billion purchase, and Nvidia says Groq racks will be online by the end of 2026\.

## Sources

- [Story source](<https://huggingnews.com/ai/nvidia-groq-3-lpx-hits-3400-tokens-per-second-in-nebius-cloud-debut-9615e020>)
- [Story source](<https://x.com/nvidia/status/2091911896538677689>)
- [Supporting source](<https://x.com/wallstengine/status/2091903357380342066>)
- [Supporting source](<https://x.com/FirstSquawk/status/2091903682724216978>)
- [Supporting source](<https://x.com/FirstSquawk/status/2091903786520572124>)
- [Supporting source](<https://x.com/CNBC/status/2091904292915388796>)
- [Supporting source](<https://x.com/JonathanRoss321/status/2091914710325289435>)

