papersTODAY 04:00 UTC
UniRank Proposes Unified Rank Allocation for Low-Rank LLM Compression
A new arXiv paper introduces UniRank, a method for deciding how much rank to assign to each weight matrix when compressing large language models via low-rank decomposition. The authors argue that uniform or hand-tuned allocation schemes overlook differences in module importance. The work is released as a revised submission on arXiv.