papersSEP 10 04:00 UTC
Divergence-based approach proposed to evaluate fidelity loss in quantized LLMs
A new arXiv paper argues that zero-shot task accuracy is an inadequate yardstick for quantized large language models, because it relies only on argmax predictions and hides changes in output distributions. The authors introduce a divergence-based method for measuring how much behavioral fidelity is lost when models undergo aggressive post-training compression for memory-constrained edge devices.