Hot Chips 2026: Cerebras lays out the future of wafer-scale AI — Nexus system architecture triples rack-scale performance, CS-6 wafer to incorporate stacked DRA

Hot Chips 2026: Cerebras lays out the future of wafer-scale AI — Nexus system architecture triples rack-scale performance, CS-6 wafer to incorporate stacked DRA

Mounting the wafer-scale engines vertically in the backpack modules lets Cerebras do away with a PCB or substrate for the wafer to handle all its supporting infrastructure. Instead, the backpack connects the large copper busbar that delivers juice to the chip directly to its back side. This close contact is important, as it minimizes power losses that occur on the way to the chip, as happens with a BGA GPU chip mounted on a PCB module with all of its power delivery circuitry located around the die.

Cerebras translates the power saved this way directly into performance in the WS-3T. The company says the losses avoided by the Nexus backpack design allow it to deliver twice as much power to the wafer-scale engine as in past designs, which leads directly to increased clock speeds and up to twice the performance of the WS-3.

Using the same base silicon as the WS-3, each WS-3T delivers twice as many sparse FP16 petaFLOPS and twice as much memory bandwidth from its SRAM. But the WS3-T is still limited to 44GB of memory across the entire wafer, and three such wafers in a CS-4 rack only scale up to 132 GB, far less than the 20.7 TB of HBM in the Vera Rubin NVL72 system and the 31 TB of AMD’s Helios .

The company doesn’t publish dense PFLOPS figures for these engines, possibly because the dataflow architectural design of the chips is specifically built to derive advantage from sparsity in a way a traditional GPU usually isn’t.

In any event, to accommodate the larger models of today and tomorrow, Cerebras will need to scale up and out. But unlike other rack-scale systems that rely on Ethernet for scale-out, Cerebras can simply connect CS-4 systems together using the same wafer-to-wafer interconnect that connects wafer-scale engines together in the Nexus rack. The company claims 2.4 Tb/s of direct scale-up bandwidth per wafer within the rack for a total of 7.2 Tb/s of inter-chip bandwidth at 2 μs latencies.

Cerebras notes that with its architecture, only the model activations need to pass between wafer-scale engines, so the relatively low bandwidth of the direct wafer connection isn't the obstacle to scaling out the system that it might seem when evaluated against the hundreds of terabytes per second of scale-up bandwidth of a system like Vera Rubin NVL72 or AMD's Helios. (The on-die fabric of the WSE-3T boasts 53.4 PB/s of bandwidth, regardless.)

The CS-4 system architecture lays the groundwork for the next-generation CS-5 accelerator, which will use new WSE silicon in 2027. For smaller models, the company says the next-generation WSE will deliver up to 10,000 tokens per second per user, while larger frontier models from labs like DeepSeek or OpenAI could run at 5,000 tokens per second per user.

As Nvidia CEO Jensen Huang has said, AI agents are impatient, and the ability to provide such vast numbers of tokens per second using specialized accelerators like the CS-4 will likely continue to be an important niche for Cerebras to exploit alongside its partners at OpenAI and AMD going forward.

(Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) (Image credit: Cerebras) Image 1 of 42 View Original

TOPICS Cerebras Jeffrey Kampman Senior Analyst, Graphics As the Senior Analyst, Graphics at Tom's Hardware, Jeff Kampman covers everything that has to do with graphics cards, gaming performance, and more. From integrated graphics processors to discrete graphics cards to the hyperscale installations powering our AI future, if it's got a GPU in it, Jeff is on it.

Key considerations

  • Investor positioning can change fast
  • Volatility remains possible near catalysts
  • Macro rates and liquidity can dominate flows

Reference reading

More on this site

Informational only. No financial advice. Do your own research.

Leave a Comment