papersSEP 10 04:00 UTC
Study Reveals Position Bias in Rubric-Based LLM-as-a-Judge Evaluations
A new arXiv paper examines large language models acting as evaluators under rubric-based protocols, a setting that has received less attention than pointwise and pairwise comparison methods. The authors find that the ordering of responses systematically influences the judge's verdicts, exposing position bias in this evaluation setup. The results suggest that pipelines relying on LLM judges may need safeguards or reordering strategies to produce reliable assessments.