Hot Chips 2026: Arm details AGI server CPU with two 70-core N3P chiplets — touts 2 TB/s UCIe fabric link and 12-channel memory controller

Hot Chips 2026: Arm details AGI server CPU with two 70-core N3P chiplets — touts 2 TB/s UCIe fabric link and 12-channel memory controller

The important point is that CMN-S3 is not just an internal CPU mesh, as Arm designed the coherent system to extend outside of the die to extend coherency beyond the die and the socket. The approach is conceptually closer to Intel's distributed Xeon 2D mesh (though Xeon is moving on to a 3D mesh with Diamond Rapids ) than AMD's EPYC architecture, where compute chiplets connect to a central I/O die that hosts the memory controllers and Infinity Fabric infrastructure. This essentially proves that Arm appears to have optimized AGI's chiplets for memory locality, bandwidth, and latency, but not exactly for compute performance density, modularity, yield, and ease of manufacturing like AMD.

Arm revealed at Hot Chips that each chiplet physically contains 70 Neoverse V3 cores, but the complete product exposes up to 136 cores, which means that four cores are redundant and are incorporated to increase yield.

Arm positions its AGI CPU primarily for AI servers and agentic AI systems, in particular. Since memory performance plays a big role in many agentic AI workloads, Arm implemented a capable coherent NUMA memory subsystem. The NUMA subsystem features two six-channel DDR5 subsystems located in each chiplet, which can potentially provide a total of up to 845 GB/s of bandwidth. If a core needs memory attached to the other chiplet, the request can cross the coherent die-to-die connection, though at a cost of latency. Arm's goal is to provide as much bandwidth per core as possible, which is why AGI supports everything up to DDR5-8800.

The DDR5 controllers within Arm's AGI CPU are quite sophisticated too. They support numerous features to maximize performance in real-world workloads, including fully out-of-order command scheduling, bank-parallelism-optimized address mapping, and programmable page policies to improve DRAM utilization and extract more effective bandwidth from the memory subsystem, while anti-starvation mechanisms help maintain predictable service under heavy load.

In addition, Arm also implements memory-bandwidth limiting and monitoring through Memory Partitioning and Monitoring (MPAM) along with QoS-based traffic prioritization and congestion feedback to manage contention when multiple cores and I/O devices compete for DRAM bandwidth. The memory subsystem also features extensive RAS capabilities, including single-DRAM-device failure correction with Chipkill-class protection, memory scrubbing, row-hammer mitigation, repair support, error injection, and RAS error logging.

Now that Arm has shared so many details about its AGI CPU, the lingering question is the performance of the processor itself. Arm still has not published conventional benchmark results such as SPEC CPU2017, SPECrate, integer/floating-point throughput, or direct socket-to-socket comparisons against current AMD EPYC or Intel Xeon processors in real-world server workloads.

The main performance claim that Arm has made is '2X performance per rack versus the latest x86 platforms ' based on estimates, which is not even remotely a detailed performance claim. Perhaps, following Nvidia's lead, Arm prefers to compare the per-rack performance of its CPUs, as they are made to work in racks. However, this is clearly an unconventional way to evaluate processors.

(Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) (Image credit: Arm) Image 1 of 21 View Original

TOPICS arm See all comments (0) Anton Shilov Social Links Navigation Contributing Writer Anton Shilov is a contributing writer at Tom’s Hardware. Over the past couple of decades, he has covered everything from CPUs and GPUs to supercomputers and from modern process technologies and latest fab tools to high-tech industry trends.

Key considerations

  • Investor positioning can change fast
  • Volatility remains possible near catalysts
  • Macro rates and liquidity can dominate flows

Reference reading

More on this site

Informational only. No financial advice. Do your own research.

Leave a Comment