papersTODAY 04:00 UTC
arXiv Paper Studies Steerability Signatures in Language Model Activations
A new arXiv preprint investigates why steering language models with contrastive representation pairs works well for some behaviors but not others. The authors look for measurable signatures in activation space that indicate how steerable a given model behavior is. The work aims to make activation-based control more predictable rather than relying on trial and error.