Abstract
We study answer readout in a 7B vision-language model adapted with DoRA on the official Med-CMR training split. The submitted system used length-normalised isolated option-text likelihood for multiple-choice questions. In a post-submission analysis on 325 held-out items from the same distribution, switching to joint option presentation and next-token option-symbol scoring increased accuracy from 0.643 to 0.954 without changing weights. This contrast bundles option presentation, prediction target, and scoring; it is not a component ablation. A model-free shortest-option baseline achieved 0.899. The untuned base model achieved approximately 0.33 under symbol scoring; where the shortest baseline failed, the adapted model answered 24/33 correctly. Thus, adaptation and more than option brevity are involved, but visual grounding is not isolated from text cues. A margin-gated hybrid reached 0.960 in-sample but averaged 0.954 under repeated five-fold cross-validation, providing no reliable advantage over symbol scoring. The analysis identifies inference configuration as a material evaluation and explains a limitation of our submitted system; it is not an official final challenge result.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/MedReason_002.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: https://openreview.net/forum?id=tBKiLfxT0f
BibTex
@InProceedings{ModUra_Answer_MICCAISAT2026,
author = { Modi, Ura AND Raval, Mehul S.},
title = { { Answer Readout as a Major Confounder in Medical Multiple-Choice Visual Question Answering } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17262},
month = {pending},
page = {pending}
}
