papersSEP 11 04:00 UTC
Study Finds LLM Simulators Can Circumvent Automated Explanation Tests
A new arXiv paper examines automated simulatability, a protocol that scores explanations by how well they let a user predict a model's outputs without relying on costly human evaluation. The authors report that when LLMs stand in for human explainees, they can bypass the explanations themselves, undermining the validity of the metric. The work suggests automated simulatability may overstate how useful an explanation really is.