papersTODAY 04:00 UTC
arXiv Paper Audits Commonsense Reasoning Benchmarks Used for LLM Evaluation
A new arXiv paper argues that commonsense reasoning in language models is typically measured with multiple-choice benchmarks such as HellaSwag and PIQA, yet these benchmarks themselves are rarely scrutinized. The authors propose a more comprehensive approach to evaluating the benchmarks, questioning how well they actually capture the capability they claim to test. The work is framed as a meta-evaluation of standard commonsense reasoning tests.