papersTODAY 04:00 UTC
Study Decomposes Transformer Representation Updates into Parallel and Perpendicular Parts
A new arXiv paper analyzes how representations inside transformer models change across layers, treating each learned update as a combination of a component that keeps the existing direction and one that shifts it elsewhere. The authors frame this as a functional geometry, aiming to explain what the model preserves versus reorients as information flows through the network.