Benchmark Tests Whether LLMs Recover Helpfulness When Users Clarify Intent
A research paper introduces CarryOnBench, a benchmark for measuring how well language models regain usefulness in multi-turn conversations after a benign user clarifies what they actually want. The authors argue that existing safety alignment work focuses on resisting adversarial prompts but largely ignores whether models can recover helpfulness in legitimate follow-ups. The benchmark targets interactive multi-turn settings rather than single-turn exchanges.