
Once the calculation is done, a PIMX_WR command transfers the output data back to the DRAM bank. Samsung noted there doesn't need to be a 1:1 relationship between reads and writes, but it's useful for this example. With a VRF size of 1 kbit, a maximum of four calculations can be written to the VRF.
With the output written back to memory banks, the host just needs to read the data from memory. The host switches to single-bank (conventional DRAM) mode and executes 16 reads to gather the output from all of the memory banks.
Samsung's LPDDR5X-PIM looks a lot like LPDDR5X. It uses a standard 561-ball array for packaging, just like LPDDR5X, and Samsung uses two 64-bit ranks with 16 GB modules. The critical number here is bandwidth. With LPDDR5X-9600, peak bandwidth is 76.8 GB/s, but that's increased by eightfold with PIM to 614 GB/s by reducing data movement and keeping basic logic local.
Samsung uses four dies per rank, for a total of eight dies. Not the various registers above, as well, as they're important for the illustration of data flow through Samsung's LPDDR5X-PIM memory.
In Samsung's preliminary benchmarks , LPDDR5X-PIM is impressive. In model run time, Samsung say a 2.28x improvement with PIM, and in tokens per second (TPS), PIM offered a 3.01x increase in performance.
The slide above shows what happened behind the scenes to gather these numbers, with Samsung using an edge AI accelerator — we're not sure which, but perhaps an early Gaia SoC — and testing Llama 3.1 with 8 billion parameters. Notably, the output is different, which one attendee pressed Samsung about. The company says optimizations are ongoing to improve accuracy, but it expects the performance benefit to remain the same.
One of the main advantages of LPDDR5X is right there in the name: low power. With PIM, power consumption becomes more of a concern, but Samsung says it doesn't expect higher power consumption overall compared to conventional DRAM. The presenter noted that peak power consumption will be "much higher" due to the bursty power draw of the PIM, but Samsung still expects overall power draw to be lower than conventional DRAM.
That comes down to extra reads/writes. Although PIM represents a power increase, decreasing the number of times data needs to move between DRAM and the host will lead to overall lower power consumption. "We're not having significant power increase," as Samsung's Karam Hwang put it.
LPDDR5X has, until recently, only had applications in consumer products. However, SOCAMM2 serviceable modules allowed Nvidia to use LPDDR5X as the memory of choice with its Vera CPU. And Intel uses LPDDR5X with its new Crescent Island AI accelerator .
Even with PIM, LPDDR5X doesn't come remotely close to the bandwidth with available with HBM, but it has a lot of applications elsewhere. Samsung's targets of server, client, and mobile are telling, with LPDDR5X-PIM accelerating edge AI on mobile and client devices, as well as arriving in lower-scope accelerators like Crescent Island.
(Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) (Image credit: Samsung) Image 1 of 19 View Original
TOPICS Samsung See all comments (0) Jake Roach Social Links Navigation Senior Analyst, CPUs Jake Roach is the Senior CPU Analyst at Tom’s Hardware, writing reviews, news, and features about the latest consumer and workstation processors.
Key considerations
- Investor positioning can change fast
- Volatility remains possible near catalysts
- Macro rates and liquidity can dominate flows
Reference reading
- https://www.tomshardware.com/pc-components/dram/SPONSORED_LINK_URL
- https://www.tomshardware.com/pc-components/dram/hot-chips-2026-samsung-makes-lpddr5x-smart-with-logic-unit-in-memory-lpddr5x-pim-is-3-01x-faster-than-lpddr5x-in-ai-inference-with-8x-the-bandwidth#main
- https://www.tomshardware.com/my-account
- Enthusiast turns a Lenovo Yoga laptop, an M.2 slot, and AMD Radeon RX 7900 XT into 'the world's stupidest' desktop for local AI chatbots — M.2 franken-rig cripp
- Trump administration weighs expanding chip tariffs to laptops, consoles, and servers, report claims — January's data center exemptions may be scrapped
- Ingenious indie hacker funds $3,000 MacBook purchase by selling advertising space on the lid — sticker space auction has already raised 111% of the price of the
- Class Is in Session: GeForce NOW Levels Up Linux, Chromebooks and More
- Hardware modder builds overclocked PS3 ‘Pro’ v3 with 3D-printed cooling mod — dual Noctua fans and server heatsinks provide FPS boost in 'almost every game'
Informational only. No financial advice. Do your own research.