papersTODAY 04:00 UTC
Value-Guided Preference Distillation Proposed for Long-Horizon Dialogue Alignment
A new arXiv paper argues that aligning multi-turn dialogue agents by matching turn-level human preferences is a poor proxy for long-term outcomes and is vulnerable to reward hacking. The authors recast long-horizon dialogue optimization as a multi-objective problem and propose distilling dense behavioral signals into value guidance for preference-based training. The method is presented as a way to optimize sparse end goals more reliably without directly optimizing them.