Amazon Web Services (AWS) has partnered with Cerebras Systems to deliver an AI inference solution that supports generative AI applications and LLM workloads.

The financial terms of the agreement have not been disclosed.

Cerebras WS-3
– Cerebras

In a statement, the companies said the solution would combine AWS Trainium-powered servers with Cerebras' wafer-scale CS-3 systems and Elastic Fabric Adapter (EFA) networking, set to be deployed on Amazon Bedrock in AWS data centers.

AWS said it would also offer open-source LLMs and Amazon Nova, the company’s own foundation models, using Cerebras hardware “later this year.”

The combined Trainium/CS-3 solution will enable “inference disaggregation,” the statement continued, an architecture which splits AI inference into two phases: a compute intensive prompt processing, or ‘prefill,’ stage, and the memory bandwidth-intensive output generation stage, known as ‘decode.’

This technique improves throughout and allows each phase to benefit from the compute architecture that best supports its needs. In this case, Trainium chips will be optimized for prefill, while the Cerebras CS-3 hardware will be optimized for decode.

“Inference is where AI delivers real value to customers, but speed remains a critical bottleneck for demanding workloads like real-time coding assistance and interactive applications,” said David Brown, VP, compute & ML services, AWS.

“What we're building with Cerebras solves that: by splitting the inference workload across Trainium and CS-3, and connecting them with Amazon’s Elastic Fabric Adapter, each system does what it's best at. The result will be inference that's an order of magnitude faster and higher performance than what's available today."

Andrew Feldman, founder and CEO of Cerebras Systems, added: “Partnering with AWS to build a disaggregated inference solution will bring the fastest inference to a global customer base.

“Every enterprise around the world will be able to benefit from blisteringly fast inference within their existing AWS environment.”

First announced in December 2020, AWS’ Trainium chips are purpose-built for ‘high-performance ML training applications in the cloud.’ The company unveiled its Trainium3 chips at its AWS Re:Invent 2025 event in December 2025, going on to sign a deal with OpenAI for the deployment of 2GW hardware in late February of this year. OpenAI has committed to use both the current Trainium3 chip and the next-generation Trainium4, currently planned for 2027.

Cerebras signed its own agreement with OpenAI this year, with the $10 billion deal expected to result in the deployment of 750MW of compute power for the generative AI giant, delivered through 2028. Last month, the chip company announced it had raised $1 billion in a Series H funding round, valuing the company at $23bn, ahead of its planned IPO later this year.