papersSEP 12 04:00 UTC
arXiv paper proposes unified per-token gating family for on-policy distillation
A new arXiv preprint introduces a family of per-token gating methods for on-policy knowledge distillation that mixes forward and reverse KL losses. The authors argue that prior approaches such as EOPD and ToDi each rely on a single fixed gating signal, and their framework generalizes these with multi-channel and bias coefficients. The work is a methodological contribution aimed at improving how distillation losses are weighted per token during training.