papersSEP 10 04:00 UTC
FlowCPO paper unifies online RL and offline preference alignment for flow and diffusion models
A new arXiv preprint presents FlowCPO, a framework that interprets preference alignment for flow and diffusion models through a single divergence-based lens. The work connects online reinforcement learning techniques with offline preference optimization, approaches that had previously been studied in isolation. The authors also examine limitations of existing forward-process alignment methods within this unified view.