papersSEP 10 04:00 UTC
Decomposing LLM-Judge Uncertainty to Target Expert Labels
A research paper addresses how to decide which LLM-judged outputs actually need human expert review. It separates the judge's uncertainty into aleatoric uncertainty, which reflects genuine disagreement among experts and cannot be reduced by more labels, and epistemic uncertainty, which signals where expert annotation would help. The goal is to spend limited expert labeling effort on the cases where it is most useful.