papersSEP 11 04:00 UTC
Study Measures Reliability of Automated Jailbreak Evaluators
A new arXiv paper examines how well automated evaluators judge whether jailbreak attacks on language models succeed, noting that human expert review is expensive and hard to scale. The authors argue that jailbreak research often fails to properly validate the evaluators it depends on, and they empirically measure those evaluators' behavior.