papersTODAY 04:00 UTC
arXiv paper critiques on-policy self-distillation, proposes RL contrastive method
A revised arXiv paper examines on-policy self-distillation, a technique that gives reasoning models dense token-level feedback by matching their output distribution to one produced with extra context such as a verified solution. The authors argue this approach yields a flawed distribution, and they introduce RLCSD, which combines reinforcement learning with contrastive self-distillation on policy. The work is a preprint and has not been peer reviewed.
arXivRLCSDcontrastive-self-distillationon-policy-self-distillationreasoning modelsreinforcement-learning
COVERAGE · 2 REPORTS · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.CLRLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation ↗TODAY 04:00 UTC
arXiv cs.LGRLCSD: Reinforcement Learning with Contrastive On-Policy Self-Distillation ↗TODAY 04:00 UTC