papersSEP 12 04:00 UTC
Study Maps Scaling Laws Behind Grokking's Delayed Generalization
A new arXiv preprint examines grokking, the phenomenon where neural networks keep memorizing training data before abruptly improving on held-out data. While prior work has focused on why this delay happens, the paper targets its quantitative structure, describing scaling laws and a phase structure that predict when the shift occurs. The authors present an arXiv preprint; the abstract excerpt provided does not detail the full experimental setup or results.