Researchers propose a method to preserve long-tailed expert knowledge in MoE fine-tuning
A new arXiv paper tackles a weakness in adapting Mixture-of-Experts models: routing layers can destabilise during supervised fine-tuning, causing rarely used experts to lose their specialised knowledge. The authors introduce a tuning approach designed to retain this long-tailed expert information and compare it with earlier anti-collapse techniques such as DenseMixer and ESFT. The work addresses a practical bottleneck for teams adapting large MoE models to downstream tasks.