papersSEP 10 04:00 UTC
Single-Direction Attack Strips Refusal Behavior From a 320B MoE Model
A new arXiv paper shows that removing one internally represented direction associated with refusals can disable a 320-billion-parameter mixture-of-experts model's ability to decline harmful requests. The technique, known as directional ablation, requires no gradient-based training or optimization—only a small set of contrastive examples to locate the direction. The authors argue this reveals that safety training in very large models may depend on a surprisingly brittle, low-dimensional mechanism.