Paper Finds Filter Metric Mismatch in Group-Relative RL With Shaped Rewards
A new arXiv paper examines how group-relative policy optimization methods such as GRPO drop rollout groups that show no contrast, through a dynamic sampling step, while real implementations let users configure the filter metric. The authors identify and measure a mismatch between that metric and the predicate used to decide which groups to discard, describing resulting "phantom advantages" when rewards are shaped. The work argues the filter metric choice is safety-critical rather than a minor implementation detail.