On October 1, 2026, Cerebras Systems reported a 5x increase in inference throughput in early tests using disaggregation, with the same number of Cerebras systems and no loss in token generation speed. The disclosure appeared in a company blog post by Isaac Tai and Zhenwei Gao titled Disaggregated Inference From the Ground Up, which opens a planned series. For businesses running large language models, the claim matters because capacity growth came without adding wafer-scale hardware.

Cerebras claims 5x inference throughput with disaggregated design

Heterogeneous disaggregation in early Cerebras tests

The company said a traditional aggregated setup required more hardware to raise capacity, while the new approach uses partner accelerators for prompt processing alongside the existing WSE footprint. A diagram in the post shows AWS Trainium and AMD Helios Instinct GPU systems as prefill options feeding a Cerebras decode pool. Cerebras said it has announced partnerships with multiple hardware partners to bring more ultrafast tokens to market. The post frames the result as early testing rather than a finished production benchmark.

Disaggregation splits inference into separate hardware pools for prefill and decode, allowing each stage to be sized, batched, and scheduled independently. After prefill builds the KV cache, that request-specific state moves to the decode pool and is loaded into memory before generation continues, while model weights remain loaded in both pools. Operators can reserve more capacity for prefill when time-to-first-token targets are strict, or give decode a larger pool for smooth streaming. The post says the handoff adds network and coordination overhead, and either pool can sit idle if capacity does not match demand.

The explanation starts from arithmetic intensity, defined as floating-point operations divided by bytes moved between memory and compute units. Adding two matrices performs 1 FLOP per 6 bytes moved, or 0.167 FLOP per byte, with the ratio constant as matrices grow, while matrix multiplication reuses loaded values across more outputs as inputs grow. During prefill, prompts of thousands or hundreds of thousands of tokens run as large parallel matrix operations between fixed weights and input tokens. During decode, tokens emerge one by one with reuse of keys and values stored in the KV cache, yet weights of hundreds of gigabytes or terabytes still move for every token.

What split prefill and decode mean for AI users

For companies operating chatbots, copilots, and document processing, separate pools make it possible to protect active responses from stalls caused by compute-intensive prefill on shared hardware. Schedulers no longer have to trade time-to-first-token against streaming smoothness and total throughput inside one queue. Batching still lets concurrent requests share weight reads, but larger batches can lengthen each decode step, so throughput can rise while per-user speed falls. Independent pools let a business prioritize fast first answers for support dialogs or steady output for long generation.

Agentic applications look like the strongest near-term beneficiary because long multi-turn workflows accumulate context across model calls and delays compound at each step. A small company buying inference by API can compare vendors on output tokens per second for long inputs, such as the cited 1,669 for Cerebras on GPT-oss-120B with 10,000 input tokens versus 708 for SambaNova and 475 for Groq. A large operator with its own serving fleet faces a different decision around idle capacity, cache transfer between colocated machines or regions, and separate scaling policies. In both cases, the published 5x figure does not by itself promise lower cost per token.

The next installments are expected to cover hardware and software stacks and economic trade-offs of disaggregated inference at scale. A useful confirmation marker will be concrete deployment detail: named partner configurations, measured transfer overhead, and pricing for split prefill and decode capacity. If those disclosures show sustained throughput gains without slower token streaming, heterogeneous inference pools will become a practical procurement option for high-volume AI workloads.