papersSEP 11 04:00 UTC
arXiv paper studies Fisher-Rao gradient flows of linear programs and natural policy gradients
A revised arXiv preprint analyzes natural gradient methods built on the Fisher information matrix of the state-action distribution. The authors connect these methods to Fisher-Rao gradient flows of linear programs, aiming to clarify why Kakade-style natural policy gradient updates converge, including in regularized settings. The work is theoretical and sits at the intersection of optimization and reinforcement learning.