Turning 99% unstructured sparsity into a wall-clock speedup
A new sparsity approach claims to overcome GPU-unfriendly scattered memory reads, converting extreme FLOP reduction into more than 20% actual speedup
1/ Unstructured sparsity in LLMs is a famous trap. You drop 99% of the FLOPs, but wall-clock time goes UP because GPUs hate scattered memory reads. A new paper finally breaks this paradox, turning 99% sparsity into a >20% actual speedup.