papersTODAY 04:00 UTC
Study finds RL training for LLMs helps easy problems more than hard ones
An arXiv paper reports that reinforcement learning does not lift large language model performance evenly across a dataset. Gains are large on problems the model can already solve and much smaller on difficult ones, a pattern the authors call the Matthew Effect. The finding suggests current RL training methods may widen the gap between easy and hard tasks.