papersSEP 10 04:00 UTC
Study Examines How Preventative Steering Defenses Hold Up During Adversarial Fine-Tuning
New arXiv research explores how language models resist harmful behavior shifts caused by malicious fine-tuning. The work evaluates preventative steering, a training-time method that injects undesirable persona vectors during fine-tuning and removes them at inference, and analyzes how the defense's effectiveness changes over the course of training. The findings suggest these safeguards require active adjustment across training phases rather than a fixed configuration.