
AMD’s 256-core Epyc 9996 ‘Venice’ claims up to a 3.4x jump over Intel Xeon competition, 20% over Nvidia Vera
(Image credit: AMD) (Image credit: AMD) (Image credit: AMD) (Image credit: AMD) Image 1 of 4 View Original
AMD's main server CPU claim at the event is that sixth-gen EPYC enables the most agents per watt, per dollar, and per rack. Endnote 9xx6-012 in the launch release states that agent counts are estimates derived from available CPU thread resources used as a proxy under a consistent theoretical workload, and that real capacity varies with workload, model, memory, software, orchestration, and system configuration. The per-rack comparison behind it is core count at a 100 kW rack power envelope, pitting the 256-core EPYC 9996 against an 88-core Nvidia Vera, AMD's own 192-core EPYC 9965, and Intel's 128-core Xeon 6980P. The per-dollar metric is based on top-of-stack thread count divided by the 1,000-unit list pricing.
The per-watt comparison in that endnote lists Nvidia Vera at 450W and Arm's AGI CPU at 300W with one thread per core, alongside Intel's Xeon 6980P at 500W and AMD's EPYC 9965 at 500W. AMD had already claimed a 3.3 times rack-level advantage over Vera in June. Mercury Research put AMD at a record 46.2% of x86 server CPU revenue in Q1 2026 , against 33.2% of units, and Arm-based designs took roughly 17.7% of server shipments in the same quarter, so the widening comparison shows where these units are going.
Starting with sixth-gen EPYC, AMD has replaced TDP with a figure it calls Default CPU Power, defined as total power consumed across the processor's compute and I/O dies at a stated performance target. AMD says both references can serve for product comparison and performance-per-watt analysis, and the endnote itself mixes the two conventions, quoting the EPYC 9956 at 400W Default CPU Power against TDP figures for the Nvidia, Intel, and Arm parts.
Helios racks pair 72 Instinct MI455X GPUs with 18 Venice CPUs, 31TB of HBM4, and 1.4 PB/s of aggregate memory bandwidth, and are in production now. AMD claims up to 30% more inference tokens per dollar than Nvidia's Vera Rubin NVL72, based on AMD Performance Labs estimates from July 2026 using a Kimi K2 Thinking workload at 32K input and 8K output, with hourly GPU pricing projections. The 34-times token throughput gain AMD quotes for MI455X over MI355X comes from AMD's own measurements on DeepSeek V4 Flash at FP4. Both, however, are vendor-provided benchmarks with no independent verification yet.
The forward roadmap runs MI500 Series GPUs in 2027 inside a Helios 500 rack built on EPYC "Verano" and Pensando "Como" and "Monza" networking, MI600 Series in 2028 inside Helios 600 on Ferrara, and Ravenna on Zen 8 in 2030.
OpenAI expects to bring Helios online from the fourth quarter of 2026, with deployments accelerating through 2027, while Meta is validating sixth-gen EPYC platforms in its labs and has begun testing Helios racks. Anthropic committed the day before the keynote to up to 2GW of MI455X GPUs in Helios systems , with the first gigawatt due in the first half of 2027. SemiAnalysis reported in February that manufacturing delays would push mass production and first production tokens on an MI455X UALoE72 system to Q2 2027; AMD software chief Anush Elangovan publicly rejected that assessment and said Helios remained on target for 2H 2026.
AMD's cautionary statement in the launch release lists the availability of essential components, naming memory supply specifically, among the risk factors that could cause results to differ from its projections. A Helios rack carries 31 TB of HBM4, and DRAM contract prices roughly doubled quarter-on-quarter in Q1 2026 before rising again in Q2.
Luke James is a freelance writer and journalist.\u00a0 Although his background is in legal, he has a personal interest in all things tech, especially hardware and microelectronics, and anything regulatory.\u00a0 ","collapsible":{"enabled":true,"maxHeight":250,"readMoreText":"Read more","readLessText":"Read less"}}), "https://slice.vanilla.futurecdn.net/13-4-25/js/authorBio.js"); } else { console.error('%c FTE ','background: #9306F9; color: #ffffff','no lazy slice hydration function available'); } Luke James Social Links Navigation Contributor Luke James is a freelance writer and journalist. Although his background is in legal, he has a personal interest in all things tech, especially hardware and microelectronics, and anything regulatory.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/pc-components/cpus/SPONSORED_LINK_URL
- https://www.tomshardware.com/pc-components/cpus/amd-splits-zen-7-into-three-epyc-families-for-2028-and-starts-selling-server-cpus-by-the-agent#main
- https://www.tomshardware.com/my-account
- 3D-printed F-14 Tomcat uses an FPGA recreation of the ‘world’s first microprocessor' — CADC’s MP944 chip controls the fighter’s swing-wing system, among other t
- 45% off: Slashed to just $1,099, this bargain RTX 5060-powered gaming laptop is discounted by $900 at HP — near half-price HyperX Omen 16 is the perfect gaming
- Geekom A9 Max 2026 review: Gorgon Point in a compact Mini PC
- AOC U32G4 32-inch 4K Dual-Refresh gaming monitor review: Solid performance and value
- At SIGGRAPH, NVIDIA Advances Graphics and Simulation With Agentic and Physical AI
Informational only. No financial advice. Do your own research.