papersSEP 10 04:00 UTC
Direct Diversity Optimization for Diverse Successful Trajectories in Preference Post-Training
A new arXiv study addresses a limitation of post-training LLM agents with trajectory-level outcome labels: such supervision offers little signal for keeping multiple distinct successful strategies that branch from the same decision state. The authors frame this as a successful trajectory diversity problem and introduce Direct Diversity Optimization, a method intended to preserve varied winning paths during preference-based post-training.