Cerebras claims 5x inference throughput gain via disaggregation
Cerebras Systems said on October 1, 2026 that disaggregation — splitting prefill and decode into separate hardware pools — raised inference throughput 5x on the same number of Cerebras systems with no loss in token generation speed. The company is developing heterogeneous disaggregation with partner accelerators such as AWS Trainium and AMD Instinct.
- 5x throughput gain on the same Cerebras WSE footprint
- Prefill and decode run in separate hardware pools
- Prefill pool can use AWS Trainium and AMD Instinct GPUs
- Artificial Analysis lists Cerebras at 1,669 output tokens/s on GPT-oss-120B
Read next
AI