papersSEP 10 04:00 UTC
New arXiv paper proposes tail-likelihood approach to reinforcement learning
A research paper on arXiv introduces a reinforcement learning method that goes beyond optimizing average reward. The authors argue that mean-based objectives can obscure differences between generative policies with equal averages but different probabilities of producing rare, high-reward outputs. The work focuses on optimizing the likelihood of these tail outcomes instead.