AMD’s upcoming RDNA 5 GPUs might improve dual-issue execution & use shader units more efficiently — LLVM patch adds new FMA instruction to ease compiling
bit_user I’d just point out a couple of things: Needing to handle VOPD3 adds more complexity to the decoders, which might not have been possible in RDNA 3 & 4, without a penalty of some sort (pipeline stages, clock speed…









