papersSEP 10 04:00 UTC
Cipher-based jailbreak attacks on LLMs work without fine-tuning, preprint claims
A preprint on arXiv explores jailbreak attacks that disguise harmful requests by encoding them with ciphers. According to the authors, these attacks can bypass a model's safety training even when the cipher is arbitrary and no fine-tuning of the target model is involved. The finding suggests defenses cannot simply rely on models being unfamiliar with a particular encoding scheme.