
When you purchase through links on our site, we may earn an affiliate commission. Here’s how it works .
AMD and Cerebras expect the new inter-rack-scale platform — based on AMD Helios rack with EPYC CPUs and Instinct MI400-series accelerators inside — to be responsible for prompt processing and large context windows, whereas Cerebras' WSE will take care of the memory-bandwidth-intensive token-generation stage.
AMD and Cerebras expect their disaggregated inference platform to deliver up to 5X higher tokens per second per watt (T/s/W) by assigning different portions of an inference workload to architectures optimized for them. Therefore, AMD Helios provides rack-scale compute capacity and large volumes of complex requests, whereas the Cerebras WSE handles latency-sensitive token generation. The two compute platforms will operate within a single inference workflow, although the companies have not disclosed additional performance data or explained how the systems will be interconnected.
The underlying idea of the AMD + Cerebras platform is essentially the same as Nvidia's CPX concept, but AMD and Cerebras assign the specialized hardware to the opposite inference stage.
Nvidia's disaggregated design separates inference into context/prefill and generation/decode. The cancelled Rubin CPX GPU with GDDR7 was optimized specifically for the compute-heavy context/prefill stage, while the regular HBM-equipped Rubin GPUs handle the memory-bandwidth-bound generation stage.
Microsoft will deploy AMD’s Helios rack-scale AI accelerator ‘at scale’ on Azure
AMD takes the wraps off its Instinct MI455X AI accelerator
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/tech-industry/artificial-intelligence/SPONSORED_LINK_URL
- https://www.tomshardware.com/tech-industry/artificial-intelligence/amd-and-cerebras-partner-on-low-latency-high-throughput-ai-inference-epyc-processors-in-helios-rack-scale-infrastructure-paired-with-cerebras-wafer-scale-engine-wse-solutions#main
- https://www.tomshardware.com/subscription
- Inside optical and the battle for scale – how the AI industry is racing to integrate photonic interconnects
- Gigabyte announces support for Chinese-made CXMT memory — pushes it to 8200 MT/s on Socket AM5
- GeForce NOW Sets Sail With ‘Path of Exile: Curse of the Allflame’ Joining the Cloud
- NVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide
- Built in Fort Worth: Wistron Opens Advanced Manufacturing Plant to Produce NVIDIA AI Systems
Informational only. No financial advice. Do your own research.