Abstract
Medical vision-language models achieve strong benchmark scores, yet it remains unclear whether these reflect genuine clinical reasoning or spurious correlations. We argue that static, correlational benchmarking is structurally insufficient for high-stakes clinical evaluation. Drawing on shortcut-learning theory and recent adversarial analyses of frontier models, we propose Clinical Adversarial Validation (CAV): clinician-guided stress testing that probes causal reliance, multimodal grounding, and reasoning stability via counterfactual perturbations, reasoning audits, and calibrated abstention. A preliminary pilot on a public chest-radiograph VQA benchmark illustrates how CAV metrics expose failures invisible to leaderboard accuracy. Adversarial, causally informed validation is necessary to bridge benchmark performance and clinical readiness.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/MedAgent_009.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: https://openreview.net/forum?id=jNcTIu9Q8O
BibTex
@InProceedings{WasAzm_Why_MICCAISAT2026,
author = { Wasi, Azmine Toushik AND Anik, Mahfuz Ahmed AND Topu, Mohsin Mahmud AND Ahsan, Md Manjurul},
title = { { Why Benchmark Accuracy Fails to Measure Clinical Reasoning in Medical Vision-Language Models: Toward Clinical Adversarial Validation } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17263},
month = {pending},
page = {pending}
}
