papersTODAY 04:00 UTC
ReCAST: Reward Credit Assignment Across Timesteps for Online Diffusion RL
A new arXiv paper introduces ReCAST, a method for assigning credit to individual timesteps when fine-tuning diffusion models with reinforcement learning from multiple reward signals. The work separates how much each reward should influence training, based on user preference, from how informative that reward actually is at a given point in the generation process. This distinction is used to drive online diffusion RL more effectively.