GAVEL: LLM Judge Protocol for Comparing Extracted Clinical Timelines Against Case Reports
Researchers introduce GAVEL, a protocol that uses a large language model as a judge to compare two clinical timelines extracted from case reports, rather than relying on a single expert reference annotation. The approach aims to address limitations in existing extraction pipelines, where evaluation is constrained by imperfect reference labels and imprecise event alignment. The work is described in a new arXiv preprint.