papersTODAY 04:00 UTC
Study Finds Rubrics Can Be Exploited to Shift LLM Judge Preferences
A new arXiv paper identifies a vulnerability in evaluation pipelines that use LLM-based judges guided by natural-language rubrics. The authors show that rubrics can serve as an attack surface, allowing subtle preference drift in judge behavior that may go unnoticed by standard benchmarks. The work highlights the need for more robust validation of rubric-driven evaluation and alignment setups.