
Hot Chips 2026: High Bandwidth Flash promises massive bandwidth and capacity, but its usability is extremely limited
Nvidia, the memory makers' largest and most leveraged customer, is reportedly testing Rubin Ultra configurations with as little as 192GB because it may not be able to source enough HBM4E. A startup with roughly $450 million raised, asking a memory maker to run a bespoke die with non-standard bank geometry on capacity that could otherwise print HBM, is negotiating from a far weaker position than that, and until the supplier is named, Raptor's 2027 volume plan rests entirely on an unknown, undisclosed dependency.
The die's 11.4 MB/mm 2 density is roughly half of HBM4's 21.9 to 26.3 MB/mm 2 , Bhoja acknowledged during the Q&A: "A lot of the drop for us was also because we used a not-so-advanced DRAM. And so if we used a more mainline DRAM, just like the HBM4 guys are doing, we would be able to push that up almost all the way to the HBM4 numbers." The density penalty therefore tracks back to whatever foundry arrangement d-Matrix currently has.
Raptor's 32GB per card stands against 192GB to 288GB for HBM4-equipped accelerators, so d-Matrix sizes deployments at rack scale instead: 72 cards carry 2.3TB, enough to hold Kimi K3 's weights at 4-bit precision with headroom for around 54 concurrent users at 1M context, by its own calculations. "Even with 32 gigabytes of memory capacity, we are able to solve SOTA models in a scale-up network, so no bits are wasted," Bhoja said.
KV cache growth works against that arithmetic over time, and co-presenter Aayush Ankit, who led Raptor's SoC architecture at d-Matrix before joining Meta's MTIA team, explained the failure mode while dismissing SRAM alternatives: "We are making this unit of compute blazingly fast. Communication becomes a bottleneck soon enough." Once models and context no longer fit a single rack, inference spills into the inter-card synchronization overhead that the vertical bandwidth was meant to eliminate, which the ISCA paper concedes.
Cerebras claimed 969 tokens per second on Llama 3.1 405B in November 2024, with third-party firm Artificial Analysis verifying the figure on live hardware, and that remains the closest published reference point for Raptor's numbers. d-Matrix's 988 tokens per second per user on a model seven times larger would be a step change if it holds water, but no third party has measured Raptor, and the comparison points in the ISCA paper are simulations anchored to early silicon characterization.
Asked by an Nvidia employee about multi-layer stacking plans, Bhoja said: "Our roadmap is still a work in progress, and we have a hard enough time trying to get one-high to work and work around all of the thermal issues of that." Samsung brought its own version of the idea to Hot Chips with zHBM, a concept that stacks HBM directly on the processor and carries a claimed 70% power-efficiency gain over an HBM4E setup, with no production timeline before HBM5. The problem for d-Matrix is that Samsung runs its own DRAM fabs; d-Matrix doesn't. Whether it can get wafers at volume remains to be seen.
(Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) (Image credit: d-Matrix) Image 1 of 30 View Original
See all comments (0) Luke James Social Links Navigation Contributor Luke James is a freelance writer and journalist. Although his background is in legal, he has a personal interest in all things tech, especially hardware and microelectronics, and anything regulatory.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/semiconductors/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/semiconductors/d-matrix-stacks-its-ai-accelerator-directly-on-custom-dram-for-100-tbs-per-card#main
- https://www.tomshardware.com/my-account
- OpenAI bans Russian ChatGPT accounts posing as a fake Israeli think tank — used VPNs to push pro-Kremlin narratives and steal academic papers for its website
- Hot Chips 2026: Micron warns HBM wafer penalty is widening with every generation — AI memory uses 3x more silicon than DDR5, company says memory wall is 'gettin
- Stock up on Seagate hard drives with $30 off the 8TB Barracuda HDD — Its lowest price at Best Buy since April
- Get 16GB of DDR5 RAM for $70 and start your AM5 gaming rig with this bundle — $568 for Ryzen 7 7700X3D, 16GB G.Skill DDR5 RAM, Asus TUF B650-E motherboard, and
- Xbox announces new disc-to-digital feature for physical media — lets players claim digital versions of their games, with support for Play Anywhere and Cloud Gam
Informational only. No financial advice. Do your own research.