Cerebras CS-4: Up to 30x Faster AI Inference, Now Shipping

Cerebras CS4

Cerebras CS-4: Up to 30x Faster AI Inference, Now Shipping

Cerebras unveils the CS-4, a rack-scale system powered by three WSE-3 Turbo wafers, delivering up to 30x faster inference than GPU systems. With a modular design that separates compute from power and cooling, it enables hyperscale deployment in hours, not days. The system achieves over 1,000 tokens per second on models exceeding 10 trillion parameters, thanks to wafer-to-wafer interconnect latency as low as 2 microseconds. First shipments begin this quarter.

By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale.

More from this day

2026-08-19