papersSEP 10 04:00 UTC
Evidence-Grounded Text Evaluation with LLM Judges Aims to Make Rubric Scoring Reliable
A research paper on arXiv introduces a method for scoring text against evaluation rubrics using large language models, addressing how black-box judge models can apply identical criteria in inconsistent ways. The approach ties each score to concrete evidence drawn from the evaluated text, making the reasoning behind judgments easier to audit and reproduce. The work is cross-listed under arXiv categories for artificial intelligence, computational linguistics, and machine learning.