papersTODAY 04:00 UTC
Paper Argues LLM-Judge Calibration in Biomedical ML Needs Four Separate Ledgers
A new arXiv paper examines how synthetic perturbations are often used as cheap calibration data for LLM evaluators in biomedical machine learning, where expert review is limited. The authors argue that a planted mutation key should not be treated as either a detector output or automatically as human ground truth. They propose formalizing four distinct ledgers to make reporting of calibration results more responsible.