Nvidia details Rubin architectural optimizations for inference – improvements target better performance and efficiency from the GPU to the rack

MoE expert weights can be distributed across GPUs in order to efficiently utilize limited per-GPU HBM capacity. Nvidia says that Rubin’s TMA has been improved to deal with the challenges of managing the growing numbers o…

Nvidia details Rubin architectural optimizations for inference – improvements target better performance and efficiency from the GPU to the rack

MoE expert weights can be distributed across GPUs in order to efficiently utilize limited per-GPU HBM capacity. Nvidia says that Rubin’s TMA has been improved to deal with the challenges of managing the growing numbers o…

Nvidia details Rubin architectural optimizations for inference – improvements target better performance and efficiency from the GPU to the rack

Rubin also improves the fundamental performance of matrix operations in the Tensor Core by doubling the amount of work those cores can perform on the K dimension, or the shared inner dimension of a pair of matrices to be…