arXiv primer surveys evaluation methods for LLMs in healthcare
A new arXiv paper reviews how large language models used in clinical and medical settings should be assessed. It argues that evaluating these systems is harder than conventional machine learning evaluation for a variety of reasons. The work is framed as an introductory guide to evaluation approaches for healthcare LLMs.