aiexpert
Home / News / Brief
Chips · Aug 21, 2026, 01:35 AM · 4 sources

Cerebras CS-4 hits 30x GPU inference speed; 4,400 tokens/sec on 120B models, first shipments Q3

Cerebras Systems unveiled the CS-4, the company's fourth-generation AI accelerator, claiming up to 30x faster inference than GPU systems. The CS-4 is built from three new Wafer Scale Engine 3 Turbo (WSE-3T) processors and represents the first member of the Cerebras Nexus rack-scale platform. On GPT-OSS-120B benchmarks, the CS-4 delivers over 4,400 tokens per second per user versus approximately 150 tokens/sec on competitive GPU systems, and is twice as fast as the predecessor CS-3.

The system integrates 44GB of SRAM on each wafer and achieves 750 petaFLOPs of compute with 129.6 petabytes per second of memory bandwidth across three wafers. A key innovation is the low-latency wafer-to-wafer interconnect, reduced to as low as 2 microseconds, supporting models with over 50 trillion parameters. The modular Nexus architecture also delivers up to 10x more throughput per watt than the CS-3, directly improving data-center economics. First shipments begin this quarter.

Cerebras positions CS-4 for disaggregated inference architectures: external prefill systems (AMD Helios, AWS Trainium) handle prompt processing, while CS-4 handles ultra-low-latency decode and token generation. This split allows agentic applications to conduct more complex reasoning and verification within the same wall-clock time. For architects, the 30x speed advantage translates directly to reduced operational latency for interactive AI and measurable cost-per-token improvements in multi-turn reasoning workloads—a direct challenge to NVIDIA's GPU dominance in inference-phase acceleration.

Sources

Everything this brief rests on
  1. 01 Primary source cerebras.ai
  2. 02 cerebras.ai cerebras.ai “Cerebras CS-4 delivers up to 30x faster AI inference than GPUs, with a modular rack-scale architecture built for hyperscale AI deployment.”
  3. 03 investors.cerebras.ai investors.cerebras.ai “The CS-4 is a rack-scale solution built from three new Wafer Scale Engines and revolutionary rack and system designs. The CS-4 is the first member of the next-generation Cerebras Nexus rack-scale platform architecture. It is up to twice as fast as the CS-3, bringing the CS-4s advantage in tokens-per-second-per-user over GPUs to up to 30x more.”
  4. 04 investors.cerebras.ai investors.cerebras.ai “In a head-to-head comparison on GPT-OSS-120B, when given identical prompts, the CS-4 delivers more than 4,400 tokens second per user (TPS/user), up to 30 times faster than GPU solutions.”