The server rack utilizes three WSE-3 Turbo chips and a new Nexus architecture to accelerate the deployment of frontier AI models. First shipments of the system start this quarter, with the company stating the device will provide a broader competitive advantage over hardware based on Nvidia chips. The hardware can run 10 trillion parameter models at 1,000 tokens per second and features 50% fewer components via modular power and I/O assemblies.

The system delivers up to 10x higher throughput per megawatt and double the overall performance of the previous CS-3 generation. Cerebras claims this hardware allows its AI accelerator to be a viable alternative to traditional GPU-based systems for high-throughput inference.

Sign in to suggest edits

Key sources

  1. SOURCE@cerebras“50% fewer components with modular compute, power, and I/O assemblies”x.com
  2. SUPPORT@techmeme“first shipments starting this quarter”x.com
  3. SUPPORT@scaling01“up to 10x higher throughput per MW”x.com
  4. SUPPORT@scaling01“10T models running at 1000 tokens/s”x.com
  5. SUPPORT@business“give it a wider advantage over Nvidia-based equipment”x.com
  6. SOURCE@andrewdfeldman“CS-4 solutions deliver both, in the same power budget.”x.com
  7. SUPPORT@dnystedt“Cerebras chips do not use HBM memory, instead relying on SRAM.”x.com
  8. SUPPORT@arjunkshah21“129.6 PB/s of memory bandwidth.”x.com
Markdown