papersYESTERDAY 16:29 UTC
Hacker News thread debates whether agreement between LLM judges signals reliability
A Hacker News discussion examines the practice of using one large language model to grade another's output, and asks whether consensus among several such judges actually indicates a correct verdict. Commenters raise concerns that models can share the same blind spots or biases, so agreement may reflect correlated error rather than genuine quality. The thread touches on how evaluation setups should be validated, for example against human raters or adversarial examples.