papersTODAY 04:00 UTC
On-Policy Self-Distillation Method Aims to Prevent Entropy Collapse in RL-Trained LLMs
A new arXiv paper proposes an approach called on-policy self-distillation that acts as a "policy reheater" for reinforcement learning with verifiable rewards. The authors target entropy collapse, a failure mode where model policies become overly concentrated, cutting rollout diversity and weakening the learning signal. The method is presented as a way to keep exploration alive during RL training of large language models.
entropy-collapseexplorationlarge-language-modelson-policy-self-distillationreinforcement-learningverifiable rewards
COVERAGE · 2 REPORTS · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.CLInternalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning ↗TODAY 04:00 UTC
arXiv cs.LGInternalize the Temperature: On-Policy Self-Distillation as Policy Reheater for Reinforcement Learning ↗TODAY 04:00 UTC