papersTODAY 04:00 UTC
K-Bench benchmark evaluates LLM unlearning in agentic settings
A new arXiv paper introduces K-Bench, a benchmark designed to test whether unlearning holds up when language models act as agents rather than just answering questions directly. The authors argue that existing benchmarks like TOFU and MUSE certify forgetting only from a model's final response, so a model that simply declines to answer is treated as having forgotten the target knowledge. They show this model-level certification does not carry over to agentic deployments, where the model's behavior unfolds over multiple steps.