papersSEP 10 04:00 UTC
Psychometric audit finds MMLU aggregate scores mainly measure factual retrieval, not reasoning
A new arXiv paper applies psychometric methods to the MMLU benchmark, analyzing how question difficulty is distributed across its aggregate score. The authors conclude that the headline number primarily reflects a model's ability to recall facts, providing limited signal about reasoning skill. The results caution against relying on MMLU alone as a measure of general AI capability.