Nvidia has moved its Groq 3 LPX low-latency inference accelerator into full production to support coding agents and multi-step reasoning workloads. Designed to extend the Vera Rubin NVL72 platform, the architecture splits inference work between Rubin GPUs for large-scale context processing and the LPX for fast token generation. In Artificial Analysis tests running Gemma 4 31B with a 100K token context, the chip reached 3,400 output tokens per second, which Nvidia says is 4x faster than the nearest alternative for latency-sensitive agentic tasks.
Nebius will be the first AI cloud to deploy the hardware through its Token Factory, with Groq also among the earliest adopters. The rollout follows a $20 billion purchase, and Nvidia says Groq racks will be online by the end of 2026.
Key sources
- SOURCE@nvidia“NVIDIA Vera Rubin NVL72 is the foundation of every AI factory”x.com
- SUPPORT@wallstengine“reached 3,400 output tokens/sec running Gemma 4 31B with a 100K-token context”x.com
- SUPPORT@firstsquawk“NVIDIA: GROQ 3 LPX HITS 3,400 OUTPUT TOKENS/SEC”x.com
- SUPPORT@firstsquawk“NVIDIA - NEBIUS FIRST AI CLOUD TO ADOPT NVIDIA GROQ 3 LPX”x.com
- SUPPORT@cnbc“Groq racks will be online this year following $20 billion purchase”x.com
- SUPPORT@jonathanross321“Groq will be among the first adopters of NVIDIA Groq 3 LPX, deploying it alongside NVIDIA Ver”x.com