Cerebras Systems unveiled the CS-4, the company's fourth-generation AI accelerator, claiming up to 30x faster inference than GPU systems. The CS-4 is built from three new Wafer Scale Engine 3 Turbo (WSE-3T) processors and represents the first member of the Cerebras Nexus rack-scale platform. On GPT-OSS-120B benchmarks, the CS-4 delivers over 4,400 tokens per second per user versus approximately 150 tokens/sec on competitive GPU systems, and is twice as fast as the predecessor CS-3.
The system integrates 44GB of SRAM on each wafer and achieves 750 petaFLOPs of compute with 129.6 petabytes per second of memory bandwidth across three wafers. A key innovation is the low-latency wafer-to-wafer interconnect, reduced to as low as 2 microseconds, supporting models with over 50 trillion parameters. The modular Nexus architecture also delivers up to 10x more throughput per watt than the CS-3, directly improving data-center economics. First shipments begin this quarter.
Cerebras positions CS-4 for disaggregated inference architectures: external prefill systems (AMD Helios, AWS Trainium) handle prompt processing, while CS-4 handles ultra-low-latency decode and token generation. This split allows agentic applications to conduct more complex reasoning and verification within the same wall-clock time. For architects, the 30x speed advantage translates directly to reduced operational latency for interactive AI and measurable cost-per-token improvements in multi-turn reasoning workloads—a direct challenge to NVIDIA's GPU dominance in inference-phase acceleration.