The human brain🧠 is incredibly efficient because it only activates the specific neurons needed for a...
By @hardmaru
David Ha (hardmaru) shares research collaboration with NVIDIA on solving the sparsity-hardware mismatch in LLMs. Their 'TwELL' format achieves >20% faster training/inference on H100 GPUs by reshaping sparsity to fit GPU architecture rather than forcing GPUs to adapt.