Paper measures and mitigates template collapse in 3D CT report generation
A new arXiv paper examines how 3D medical vision-language models can write fluent radiology-style reports while still missing critical findings and producing highly repetitive output. The authors characterize this behavior, which they call template collapse, noting that models default to generic phrasing that under-reports rare pathologies. They then propose ways to measure the problem and reduce it.