Cerebras CS-4: Up to 30x Faster AI Inference, Now Shipping

Cerebras CS4

Cerebras CS-4: Up to 30x Faster AI Inference, Now Shipping

Cerebras unveils the CS-4, a rack-scale system powered by three WSE-3 Turbo wafers, delivering up to 30x faster inference than GPU systems. With a modular design that separates compute from power and cooling, it enables hyperscale deployment in hours, not days. The system achieves over 1,000 tokens per second on models exceeding 10 trillion parameters, thanks to wafer-to-wafer interconnect latency as low as 2 microseconds. First shipments begin this quarter.

By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale.
  1. syntaxing

    I think the fun takeaway from this is that GPT 5.4 is probably 45B active parameters and GPT 5.6 Sol is closer to 50B.

  2. sreekanth850

    AMD along with cerebras may probably compete with NVIDIA monopoly in near future. Also, NVIDIA will have competition form multiple companies. Just my prediction.

  3. reilly3000

    > CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters

    Oops did they just out GPT-5.6 sol’s parameter count?

  4. ethanzhang1024

    If cerebars is performing well, why didn't its predecessor, server S-3, become the largest API token provider on OpenRouter, surpassing the official model releases?

  5. aneryu

    It would be even better if a version available to individual users were released soon.

  6. kobe_bryant

    can these vibe coded sites please set a max width and overflow so their sites work fine on mobile

  7. anonymous_user9

    Conspicuously missing: power consumption figures

  8. selimonder

    That "GPU" comparison is the vaguest i seen so far

More from this day

2026-08-19