
Mounting the wafer-scale engines vertically in the backpack modules lets Cerebras do away with a PCB or substrate for the wafer to handle all its supporting infrastructure. Instead, the backpack connects the large copper busbar that delivers juice to the chip directly to its back side. This close contact is important, as it minimizes power losses that occur on the way to the chip, as happens with a BGA GPU chip mounted on a PCB module with all of its power delivery circuitry located around the die.
Cerebras translates the power saved this way directly into performance in the WS-3T. The company says the losses avoided by the Nexus backpack design allow it to deliver twice as much power to the wafer-scale engine as in past designs, which leads directly to increased clock speeds and up to twice the performance of the WS-3.
Using the same base silicon as the WS-3, each WS-3T delivers twice as many sparse FP16 petaFLOPS and twice as much memory bandwidth from its SRAM. But the WS3-T is still limited to 44GB of memory across the entire wafer, and three such wafers in a CS-4 rack only scale up to 132 GB, far less than the 20.7 TB of HBM in the Vera Rubin NVL72 system and the 31 TB of AMD’s Helios .
The company doesn’t publish dense PFLOPS figures for these engines, possibly because the dataflow architectural design of the chips is specifically built to derive advantage from sparsity in a way a traditional GPU usually isn’t.
In any event, to accommodate the larger models of today and tomorrow, Cerebras will need to scale up and out. But unlike other rack-scale systems that rely on Ethernet for scale-out, Cerebras can simply connect CS-4 systems together using the same wafer-to-wafer interconnect that connects wafer-scale engines together in the Nexus rack. The company claims 2.4 Tb/s of direct scale-up bandwidth per wafer within the rack for a total of 7.2 Tb/s of inter-chip bandwidth at 2 μs latencies.
Cerebras notes that with its architecture, only the model activations need to pass between wafer-scale engines, so the relatively low bandwidth of the direct wafer connection isn't the obstacle to scaling out the system that it might seem when evaluated against the hundreds of terabytes per second of scale-up bandwidth of a system like Vera Rubin NVL72 or AMD's Helios. (The on-die fabric of the WSE-3T boasts 53.4 PB/s of bandwidth, regardless.)
The CS-4 system architecture lays the groundwork for the next-generation CS-5 accelerator, which will use new WSE silicon in 2027. For smaller models, the company says the next-generation WSE will deliver up to 10,000 tokens per second per user, while larger frontier models from labs like DeepSeek or OpenAI could run at 5,000 tokens per second per user.
As Nvidia CEO Jensen Huang has said, AI agents are impatient, and the ability to provide such vast numbers of tokens per second using specialized accelerators like the CS-4 will likely continue to be an important niche for Cerebras to exploit alongside its partners at OpenAI and AMD going forward.
(Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) Image 1 of 42 View Original
TOPICS Cerebras Jeffrey Kampman Senior Analyst, Graphics As the Senior Analyst, Graphics at Tom's Hardware, Jeff Kampman covers everything that has to do with graphics cards, gaming performance, and more. From integrated graphics processors to discrete graphics cards to the hyperscale installations powering our AI future, if it's got a GPU in it, Jeff is on it.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/hot-chips-2026-cerebras-lays-out-the-future-of-wafer-scale-ai-nexus-system-architecture-triples-rack-scale-performance-cs-6-wafer-to-incorporate-stacked-dram#main
- https://www.tomshardware.com/my-account
- Hot Chips 2026: Fujitsu's Monaka CPU stacks its entire cache on a separate 5nm die and narrows to 256-bit SVE2 — 350W and 500W SKUs due in 2027
- Class Is in Session: GeForce NOW Levels Up Linux, Chromebooks and More
- Xbox marks 25 years with eye-catching translucent green PC accessories — Razer mouse, keyboard, and earbuds available to pre-order now
- Nvidia expects to sell $20 billion of Vera Rubin systems in Q3 as shipments begin — figure would account for 20% of its data center revenue mix, marks fastest r
- Beat PC component price rises with this unbelievably good value 9800X3D gaming desktop — get an RTX 5070, 32GB of DDR5, and a 2TB SSD all for less than $2,000
Informational only. No financial advice. Do your own research.