papersTODAY 04:00 UTC
Study questions LLM-as-a-judge validity for psychological depth evaluations
A new arXiv paper examines whether LLM judges can reliably measure psychological depth in open-ended model outputs. The authors argue that a judge's correlation with human ratings on its development set does not guarantee valid scoring when candidate responses are closely matched and human preferences are subjective. The work points to limits of LLM-as-a-judge setups that are increasingly used to evaluate generated text.
arXivLLM evaluationLLM-as-a-judgehuman preference evaluationopen-ended text generationpsychological depth evaluation
COVERAGE · 2 REPORTS · LINKS GO TO THE ORIGINAL OUTLETS
arXiv cs.CLDoes Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging ↗TODAY 04:00 UTC
arXiv cs.LGDoes Reasoning Improve Psychological Depth in Large Language Models? It Depends on Who's Judging ↗TODAY 04:00 UTC