Abstract
Multimodal large language models (LLMs) are increasingly explored for radi-ology applications, yet standard benchmarks typically report single diagnostic accuracy without assessing whether models produce consistent reasoning across repeated queries. We propose a Three-Level Hierarchical Evaluation Framework that sequentially assesses (1) anatomical location, (2) imaging fea-ture description, and (3) diagnosis, measuring both correctness and cross-repetition consistency at each level. A hierarchical gating mechanism advanc-es cases only when both criteria are satisfied, quantifying the fraction of cases achieving “reliably correct understanding.” Three state-of-the-art multimodal LLMs—Claude Opus 4.6, Gemini 3 Pro, and Kimi K2.5—were evaluated on 121 ultrasound cases under image-only, text-only, and image + text conditions, each repeated three times (N = 3,267 total responses). While ungated diagnos-tic accuracy reached up to 43.5%, the full hierarchical gate pass rate never ex-ceeded 8.3%. Within-condition similarity analysis revealed that image-only inputs yielded the lowest response reproducibility, whereas text-only inputs exhibited the highest stability, suggesting current vision encoders struggle with the inherent noise and visual complexity of ultrasound images, potentially making image inputs a primary source of response instability. These findings demonstrate that conventional metrics substantially overestimate the clinical reliability of multimodal LLMs, underscoring that hierarchical frameworks jointly assessing accuracy and reproducible reasoning must become standard in medical image interpretation benchmarking.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/ASMUS_062.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: https://openreview.net/forum?id=nDL7aVMLla
BibTex
@InProceedings{HanTae_Beyond_MICCAISAT2026,
author = { Han, Taewon AND Shin, Jaeseung},
title = { { Beyond Single Diagnostic Accuracy: Three-Level Hierarchical Evaluation of Accuracy and Reliability in Multimodal LLM for Ultrasound Interpretation } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17276},
month = {pending},
page = {pending}
}
