
A Rubin GPU carries 288GB of HBM4 , roughly 576 times the memory of a single LP30, so a 31-billion-parameter model at FP8 needs on the order of 62 LPUs to hold its weights, and a large mixture-of-experts model runs into four figures of chips across several racks. Capacity is the cost of the SRAM-only design, and it's why Nvidia is describing the LPU as for decode rather than as a general-purpose replacement for its GPUs.
Determinism lets the compiler predict power draw cycle by cycle, which Nvidia uses to pre-order current from the rack's regulators ahead of demand, cutting voltage droop by more than 60% and overshoot by more than 70% against an uncompensated load. The same per-block scheduling lets the hardware equalize heat instead of throttling to the hottest tile, which Arsovski put at roughly 10% to 11% additional performance under a fixed thermal limit. "By doing this, we can actually get more utilization of the chip under the same thermal limit, basically. So we can actually get, again, about 10 to 11% more performance under the same thermal limit. So this is another benefit of deterministic execution."
Across racks, Nvidia synchronizes chips to a single virtual clock in what it calls a plesiosynchronous network, with each chip acting as both processor and router so the fabric needs no adaptive routing or congestion sensing, and clock drift between chips is compensated at the chip-to-chip links. Asked during Q&A about the blast radius of a chip that fails mid-workload, Arsovski said users "would experience the exact same as any other hardware in the industry" and would "just checkpoint it or reconfigure the hardware."
Nvidia is pitching the LPX rack as a decode co-processor bolted onto Vera Rubin NVL72 , with Rubin GPUs handling the compute-heavy prefill phase and building the KV cache while the LPUs generate output tokens. Nvidia showed three ways to divide the work: disaggregated prefill and decode; attention-FFN disaggregation, which keeps attention and its cache on GPU HBM while the LPU runs the feed-forward layers; and external-draft speculative decoding, where a small model on the LPU proposes tokens that the GPU verifies in parallel, with only draft tokens crossing the link.
An FPGA bridges the synchronous LPU domain and the asynchronous world of host I/O and GPU hand-offs, and Nvidia's Dynamo runtime, together with an LPU extension to CUDA, orchestrates the split. The company put the gains from these modes at roughly three-to-five-times over Rubin alone on a two-trillion-parameter workload with a 400K-token cached context, all Nvidia-measured.
Cerebras used the same Hot Chips session to present its CS4 wafer-scale system, which chief system architect Jean-Philippe Fricker said runs up to 30 times faster than GPUs and doubles the token rate of the CS3 while carrying 10 times the token capacity. Each CS4 rack packs three wafer-scale engines into a new modular platform Cerebras calls Nexus, built around pluggable compute "backpacks" that separate power, compute, and I/O, and Fricker put its memory bandwidth at 43 PB/s, which he told the audience was "2,000 times higher memory bandwidth than Nvidia's next-generation Rubin chip." Cerebras also has a partner for the prefill side of the same problem: it agreed in July to pair AMD Helios GPUs for prefill with its wafer-scale engines for decode, the same division of labor Nvidia now builds in-house with Groq.
Nvidia pulled the Rubin CPX, its own GDDR7-based long-context accelerator, to focus on shipping the LPU this year, a decision VP Ian Buck laid out at GTC 2026 . The $20 billion deal that produced the LP30 was structured as a non-exclusive IP license plus the hiring of Ross, president Sunny Madra, and most of Groq's engineers, a form that avoided a formal merger review. Arsovski opened the Hot Chips talk by calling it "a pinch me moment for the Groq team that's now integrated into the Nvidia group."
Senators Elizabeth Warren and Richard Blumenthal wrote to the FTC and to Nvidia in early 2026, arguing the arrangement acquired Groq "in all but name," and no formal, deal-specific investigation has been confirmed as of late August.
(Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) (Image credit: Nvidia) Image 1 of 44 View Original
TOPICS Nvidia See all comments (0) Luke James Social Links Navigation Contributor Luke James is a freelance writer and journalist. Although his background is in legal, he has a personal interest in all things tech, especially hardware and microelectronics, and anything regulatory.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/semiconductors/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/semiconductors/nvidia-presents-groq-3-lpx-architecture-and-unveils-its-first-third-party-inference-benchmark#main
- https://www.tomshardware.com/my-account
- Nvidia custom 'NVHBM' promises 30% higher bandwidth, 15% lower power than commodity HBM4e — custom base die and PHY will be available to NVLink Fusion partners
- Bill Gates calls for some jobs to be ‘Human Reserved,’ suggests taxing AI tokens and robots — billionaire says that ‘AI era will be one of the most turbulent ti
- Save up to 61% on gaming PCs, laptops, and more in HP's Labor Day 2026 sale — huge discounts on a range of hardware, monitors, and peripherals
- Quantum computing used in 'first commercial game development' — IBM simulator generated maps, characters, and graphics in C.L.A.Y. RPG
- Hot Chips 2026: Intel dives deep on Crescent Island AI accelerator — larger caches and deeper XMX engines target maximum AI FLOPS per watt
Informational only. No financial advice. Do your own research.