papersTODAY 04:00 UTC
Researchers propose KL-projected natural policy gradient algorithms for Nash equilibrium learning in Markov potential games
A new arXiv paper studies decentralized learning of Nash equilibria in infinite-horizon discounted Markov games where agents only receive bandit feedback. The authors develop KL-projected natural policy gradient methods for both episodic and fully online asynchronous settings, aimed at Markov alpha-potential games. They also discuss applications to Markov congestion games.