arXiv paper studies how reinforcement learning reshapes LLMs using mechanistic interpretability
A new arXiv preprint examines what large language models actually learn during reinforcement learning training, approaching the question through mechanistic interpretability rather than behavior alone. The authors argue that earlier explanations of RL's effects have mostly been behavioral, and they propose using sparse autoencoders to analyze internal changes. The work is released under a fixed-SAE track.