Abstract
Automatic report generation for chest CT has advanced rapidly, yet how well pretrained models perform on real clinical data remains largely untested. We evaluate two published report generation models on an out-of-distribution clinical dataset and discuss the reproducibility challenges that frequently accompany freely available models. To enable the evaluation, we assemble and curate an institutional dataset of raw clinical data into model-ready inputs. Comparing it to a public benchmark reveals substantial differences, including more detailed localization, explicit progression descriptions, and frequent quantitative measurements in the clinical reports. Assessing the selected models with natural language generation and clinical efficacy metrics, we find that they only marginally outperform naive baselines. A hallucination analysis indicates that the evaluated models primarily reproduce training-set language while making limited use of the visual input. These findings underscore the importance of evaluating report generation models on realistic clinical data and of publishing complete reproduction information as a standard of scientific practice.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/ELAMI_017.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: Not Available
BibTex
@InProceedings{HirDom_Evaluating_MICCAISAT2026,
author = { Hirsch, Dominik AND Ulrich, Hannes AND Handels, Heinz AND Ehrhardt, Jan},
title = { { Evaluating Medical Report Generation on Real Clinical Data: A Case Study } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17262},
month = {pending},
page = {pending}
}
