papersSEP 10 04:00 UTC
XAI-Arena: Testing whether LLMs can judge the quality of explainable AI explanations
A new arXiv paper introduces XAI-Arena, a study of whether large language models can reliably evaluate explanations produced by explainable AI methods. The authors note that current evaluation relies heavily on subjective human judgment, which hurts reproducibility, scalability, and comparability across studies. The work explores automated, LLM-based assessment as a potential alternative to manual expert reviews.