papersSEP 10 04:00 UTC
Training trajectories determine circuit removability in annealable soft-prior Transformers
Researchers asked whether retrieval circuits that small Transformers learn with the help of soft positional priors keep functioning once that prior is taken away. They tested this using a model whose prior-based attention biases can be gradually annealed out during training. The results indicate that the specific training trajectory, not just the architecture, decides whether a learned circuit can stand on its own after the prior is removed.
circuit-removabilitymechanistic-interpretabilitypositional-priorsretrieval-circuitstraining dynamicstransformer-circuits
COVERAGE · 1 REPORT · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.LGTraining Trajectories Determine Circuit Removability in Annealable Soft-Prior Transformers ↗SEP 10 04:00 UTC