Refinement-Based Flow Policy Optimization for Online Reinforcement Learning
A new arXiv paper introduces Refinement-based Flow Policy Optimization, a method for using flow-based policies in online reinforcement learning. Standard flow matching needs samples from the target distribution, which is unavailable when the desired action distribution is only implicitly defined. The approach is presented as a refinement procedure that sidesteps this requirement in online RL settings.