papersSEP 10 04:00 UTC
VestigeKV paper uses NoPE-MLA's vestigial branch for KV cache compression
A new arXiv paper shows that attention-based KV cache pruning methods break down on NoPE-MLA models, with H2O and SnapKV retrieving almost none of the injected needles at 8x compression. The proposed VestigeKV method instead relies on a vestigial branch within the NoPE-MLA cache that carries its own sparse-attention signal, allowing token selection before the future queries are known.