Cerebras CS-4: Up to 30x Faster AI Inference, Now Shipping
Cerebras CS4

Cerebras unveils the CS-4, a rack-scale system powered by three WSE-3 Turbo wafers, delivering up to 30x faster inference than GPU systems. With a modular design that separates compute from power and cooling, it enables hyperscale deployment in hours, not days. The system achieves over 1,000 tokens per second on models exceeding 10 trillion parameters, thanks to wafer-to-wafer interconnect latency as low as 2 microseconds. First shipments begin this quarter.
By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale.
- syntaxing
I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.
- sreekanth850
AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.
- reilly3000
> CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters
Oops did they just out GPT-5.6 sol’s parameter count?
- ethanzhang1024
If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?
- aneryu
It would be even better if a version available to individual users were released soon.
- kobe_bryant
can these vibe coded sites please set a max width and overflow so their sites work fine on mobile
- anonymous_user9
Conspicuously missing: power consumption figures
- selimonder
That "GPU" comparison is the vaguest i seen so far