
A PIMX_RD reads the weight data from the DRAM bank, feeding into the MAC trees alongside the data from the SRF. Once the calculation is done, the output vector from each operation is written into the Vector Register File.
Once the calculation is done, a PIMX_WR command transfers the output data back to the DRAM bank. Samsung noted there doesn't need to be a 1:1 relationship between reads and writes, but it's useful for this example. With a VRF size of 1 kbit, a maximum of four calculations can be written to the VRF.
With the output written back to memory banks, the host just needs to read the data from memory. The host switches to single-bank (conventional DRAM) mode and executes 16 reads to gather the output from all of the memory banks.
Samsung's LPDDR5X-PIM looks a lot like LPDDR5X. It uses a standard 561-ball array for packaging, just like LPDDR5X, and Samsung uses two 64-bit ranks with 16 GB modules. The critical number here is bandwidth. With LPDDR5X-9600, peak bandwidth is 76.8 GB/s, but that's increased by eightfold with PIM to 614 GB/s by reducing data movement and keeping basic logic local.
Samsung uses four dies per rank, for a total of eight dies. Not the various registers above, as well, as they're important for the illustration of data flow through Samsung's LPDDR5X-PIM memory.
In Samsung's preliminary benchmarks , LPDDR5X-PIM is impressive. In model run time, Samsung say a 2.28x improvement with PIM, and in tokens per second (TPS), PIM offered a 3.01x increase in performance.
The slide above shows what happened behind the scenes to gather these numbers, with Samsung using an edge AI accelerator — we're not sure which, but perhaps an early Gaia SoC — and testing Llama 3.1 with 8 billion parameters. Notably, the output is different, which one attendee pressed Samsung about. The company says optimizations are ongoing to improve accuracy, but it expects the performance benefit to remain the same.
One of the main advantages of LPDDR5X is right there in the name: low power. With PIM, power consumption becomes more of a concern, but Samsung says it doesn't expect higher power consumption overall compared to conventional DRAM. The presenter noted that peak power consumption will be "much higher" due to the bursty power draw of the PIM, but Samsung still expects overall power draw to be lower than conventional DRAM.
That comes down to extra reads/writes. Although PIM represents a power increase, decreasing the number of times data needs to move between DRAM and the host will lead to overall lower power consumption. "We're not having significant power increase," as Samsung's Karam Hwang put it.
LPDDR5X has, until recently, only had applications in consumer products. However, SOCAMM2 serviceable modules allowed Nvidia to use LPDDR5X as the memory of choice with its Vera CPU. And Intel uses LPDDR5X with its new Crescent Island AI accelerator .
Even with PIM, LPDDR5X doesn't come remotely close to the bandwidth with available with HBM, but it has a lot of applications elsewhere. Samsung's targets of server, client, and mobile are telling, with LPDDR5X-PIM accelerating edge AI on mobile and client devices, as well as arriving in lower-scope accelerators like Crescent Island.
(Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) Image 1 of 19 View Original
TOPICS Samsung See all comments (0) Jake Roach Social Links Navigation Senior Analyst, CPUs Jake Roach is the Senior CPU Analyst at Tom’s Hardware, writing reviews, news, and features about the latest consumer and workstation processors.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/pc-components/dram/SPONSORED_LINK_URL
- https://www.tomshardware.com/pc-components/dram/hot-chips-2026-samsung-makes-lpddr5x-smart-with-logic-unit-in-memory-lpddr5x-pim-is-3-01x-faster-than-lpddr5x-in-ai-inference-with-8x-the-bandwidth#main
- https://www.tomshardware.com/my-account
- Nvidia’s GB300-powered DGX Station desktop tower listed for nearly $100,000 online — Enterprise AI powerhouse now available to buy for mere mortals with lots of
- Class Is in Session: GeForce NOW Levels Up Linux, Chromebooks and More
- 32GB of DDR5 6000 RAM now costs $400 — the latest price hike puts DIY PC building further out of reach than ever before
- Marvell VP pushes for DDR4 recycling for use in CXL memory, amid the worst DRAM shortage in years — company introduces three-tier AI memory infrastructure
- AI coder gets Doom running on a custom CPU designed by GPT-5.6 Sol — game viewport is overlaid on a pulsing schematic of the CPU in Turing Complete's sandbox en
Informational only. No financial advice. Do your own research.