Developers building agentic applications will have access to new compute capacity through a multiyear agreement between General Compute and Cerebras. The company will integrate Cerebras systems into its cloud infrastructure to provide inference that is 20x faster than current options. This deployment uses a disaggregated architecture where Nvidia GPUs manage prefill operations and Cerebras hardware handles inference, with first tokens going live in Q1 2027.

General Compute recently raised $400M in debt to finance this and other alternative AI hardware deployments. By matching specific workloads to the hardware most efficient for those tasks, the company aims to optimize the economics of its AI infrastructure and improve overall inference speed for its customers.

Sign in to suggest edits

Key sources

  1. SOURCEmarketbrief.now
  2. SOURCEhuggingnewshuggingnews.com
Markdown