papersSEP 12 04:00 UTC
Belief-Shift Branching Targets Credit Assignment in Tree-Structured RL
A new arXiv paper proposes forking rollout trees at points where the model's beliefs shift, rather than at arbitrary intermediate steps, to assign credit in critic-free reinforcement learning with verifiable rewards. Because each fork adds sampling cost, the authors argue that concentrating branches on belief changes yields step-level value estimates more efficiently. The approach is aimed at improving how tree-structured rollouts trade compute for credit assignment.