papersTODAY 04:00 UTC
Variance-Penalized Actor-Critic Method Avoids a Second Critic for Risk-Sensitive RL
Researchers propose a nonparametric approach to variance-penalized reinforcement learning that trades expected return for policy stability without training a separate variance critic. The work frames risk-sensitive RL through statistical inference, aiming to cut the extra computation and complexity that online variance estimation usually requires. It is a new arXiv preprint in machine learning.