Verifier-Gated Multi-Expert Distillation Aimed at Scientific Reasoning
A new arXiv paper examines multi-teacher on-policy distillation, the technique of training specialist models and then transferring their abilities to a single student using the student's own generated outputs. The authors propose assigning supervision token by token rather than sequence by sequence, with a verifier deciding which expert teacher should guide each token. The method is aimed at scientific reasoning tasks.