Cerebras just announced the CS-4, and the numbers demand attention — not because they're the biggest, but because Cerebras is the only wafer-scale company actually shipping product to hyperscalers.
What changed
The CS-4 is a rack-scale system built around the WSE-3 Turbo. Three chips per system, each delivering up to 2x the speed of the previous generation. The modular "backpack" design folds the wafer, power conversion, and direct liquid cooling into a self-contained assembly — no exotic plumbing needed.
The key numbers:
- 30x faster inference vs production GPU systems on the same workloads
- 1,000+ tokens/second on models exceeding 10 trillion parameters
- 2µs wafer-to-wafer interconnect latency — this is the architectural moat. Most distributed systems lose coherence at scale; Cerebras wired the chips together at the wafer level
- 10x more throughput per watt than the CS-3
These aren't paper specs. Cerebras names hyperscale deployments and benchmarks against production GPU clusters, not theoretical peak FLOPS.
Why this matters now
This lands the same week Etched's valuation hit $21B — double what it was a month ago. The inference hardware market is signaling something: GPUs were designed for training, and the market is voting with money that inference needs a different architecture.
Cerebras' bet is that the bottleneck isn't compute — it's coherence. Most distributed inference systems spend more time moving data between chips than computing on it. The wafer-scale approach eliminates that tax entirely. 2µs interconnect latency isn't an optimization; it's a category difference.
CS-4 also introduces the Nexus Rack-Scale Platform, a modular architecture that separates compute, power, and I/O into independently serviceable units. This is the part that matters for adoption: hyperscalers don't buy chips, they buy deployable systems. Nexus makes CS-4 slot into existing datacenter workflows without custom rack engineering.
Verdict: Cerebras was always the dark horse of the custom AI silicon race. CS-4 makes it a real contender for the inference tier — and Nexus makes it deployable. The wafer-scale thesis is no longer theoretical; it's shipping.