papersSEP 12 04:00 UTC
Paper compares diff-in-means and INLP for finding refusal directions in LLMs
A preprint revisits the finding that refusal behavior in safety-tuned chat models is controlled by a single linear direction in the residual stream, which can be recovered by taking the difference in means between harmful and harmless activations. The authors compare this diff-in-means approach with INLP, an iterative nullspace projection method, to see whether a one-direction account holds up. The work is presented as a preliminary comparison of the two techniques.