Researchers from MIT, NVIDIA, and Zhejiang University Propose TriAttention: A KV Cache Compression Method That Matches Full Attention at 2.5× Higher Throughput
By Asif Razzaq
Researchers from MIT, NVIDIA, and Zhejiang University propose TriAttention, a KV cache compression method that matches full attention quality while achieving 2.5× higher throughput. This directly addresses the memory bottleneck in long-chain reasoning models like DeepSeek-R1 and Qwen3.