papersSEP 10 04:00 UTC
Single-Direction Attack Strips Refusal Behavior From a 320B MoE Model
A new arXiv paper shows that removing one internally represented direction associated with refusals can disable a 320-billion-parameter mixture-of-experts model's ability to decline harmful requests. The technique, known as directional ablation, requires no gradient-based training or optimization—only a small set of contrastive examples to locate the direction. The authors argue this reveals that safety training in very large models may depend on a surprisingly brittle, low-dimensional mechanism.
COVERAGE · 3 REPORTS · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.CLHow Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE ↗SEP 10 04:00 UTC
arXiv cs.AIHow Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE ↗SEP 10 04:00 UTC
arXiv cs.LGHow Fragile Is Safety Alignment at Frontier Scale? A Single-Direction Attack on a 320B MoE ↗SEP 10 04:00 UTC