papersTODAY 04:00 UTC
Study maps how harm refusal is routed across LLM model families
A new preprint argues that measuring how easily refusal behavior can be ablated captures only part of what a model has encoded about harmful requests. The authors propose a "harm-keyed routing" account, in which refusal draws on a limited subset of the model's internal representation, and document cases where models diverge from that pattern. They test the idea across several model families to characterize when the routing holds and when it breaks down.