Abstract

Medical vision-language models achieve strong benchmark scores, yet it remains unclear whether these reflect genuine clinical reasoning or spurious correlations. We argue that static, correlational benchmarking is structurally insufficient for high-stakes clinical evaluation. Drawing on shortcut-learning theory and recent adversarial analyses of frontier models, we propose Clinical Adversarial Validation (CAV): clinician-guided stress testing that probes causal reliance, multimodal grounding, and reasoning stability via counterfactual perturbations, reasoning audits, and calibrated abstention. A preliminary pilot on a public chest-radiograph VQA benchmark illustrates how CAV metrics expose failures invisible to leaderboard accuracy. Adversarial, causally informed validation is necessary to bridge benchmark performance and clinical readiness.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/MedAgent_009.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to Open Review

Open Review Page: https://openreview.net/forum?id=jNcTIu9Q8O

BibTex

@InProceedings{WasAzm_Why_MICCAISAT2026,
        author = { Wasi, Azmine Toushik AND Anik, Mahfuz Ahmed AND Topu, Mohsin Mahmud AND Ahsan, Md Manjurul},
        title = { { Why Benchmark Accuracy Fails to Measure Clinical Reasoning in Medical Vision-Language Models: Toward Clinical Adversarial Validation } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 17263},
        month = {pending},
        page = {pending}
}


back to top