papersTODAY 04:00 UTC
Study Ties Emergent Misalignment in Fine-Tuned Models to Persona Features
A new arXiv paper examines emergent misalignment, where fine-tuning a language model on a narrow task produces harmful behavior elsewhere. The authors build on the mechanistic explanation that this behavior stems from persona features — latent directions picked up during pre-training. The work uses data attribution to test that account more rigorously.