Abstract
Deep learning models for medical imaging are increasingly assessed across multiple external datasets before clinical deployment. However, strong external validation does not necessarily demonstrate deployment readiness, particularly in underrepresented healthcare settings. Using diabetic retinopathy (DR) screening as a case study, we evaluate two publicly available models—an interpretable sparse BagNet and a conventional ResNet-50—originally trained on the Kaggle dataset and externally validated on ten publicly available retinal imaging datasets. Without retraining or adaptation, both models were assessed on a private retinal fundus dataset from the Nakuru Eye Disease Study in Kenya using the original evaluation protocol. After reproducing the published external validation results, evaluation on the Kenyan cohort revealed substantial reduced sensitivity while specificity remained high. The reduced sensitivity was entirely driven by missed Grade 1 DR cases, with no false negatives for Grades 2–4. Threshold optimization partially improved sensitivity but did not recover previous external validation performance, while qualitative analysis of the interpretable model provided complementary evidence for local model auditing. These findings demonstrate how local cross-context evaluation, operating-threshold assessment, grade-specific error analysis, and qualitative auditing can complement conventional external validation and provide context-specific evidence to inform deployment decisions in underrepresented healthcare settings.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/AFRICAI_034.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: https://papers.miccai.org/miccai-2026-sat/supp/AFRICAI_034_supp.pdf
Link to Open Review
Open Review Page: https://openreview.net/forum?id=cH95t7RWFO
BibTex
@InProceedings{DjoKer_Beyond_AFRICAI_MICCAISAT2026,
author = { Djoumessi, Kerol AND Mathenge, Ciku Wanjiku AND Berens, Philipp},
title = { { Beyond External Validation: Evaluating the Deployment Readiness of Medical Imaging AI in an African Clinical Context through Diabetic Retinopathy Screening } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17264},
month = {pending},
page = {pending}
}
