Nvidia details Rubin architectural optimizations for inference – improvements target better performance and efficiency from the GPU to the rack
MoE expert weights can be distributed across GPUs in order to efficiently utilize limited per-GPU HBM capacity. Nvidia says that Rubin’s TMA has been improved to deal with the challenges of managing the growing numbers o…




