Abstract
Optical coherence tomography (OCT) drives the management of retinal disease, yet clinical decisions hinge on the change between visits rather than the appearance of any single scan. General medical vision–language models (VLMs) reason over images in isolation and cannot ground longitudinal judgements in comparable historical trajectories. We present TrajRAG, a training-free, trajectory-aware multimodal retrieval-augmented framework for OCT change detection: each visit pair is encoded by a biomedical CLIP model into a composite transition vector, a temporal-gap-penalised, class-prior-reweighted selector with BM25 compression assembles precedent cases from a patient-disjoint memory, and a medical VLM classifies the change and forecasts the next visit, with temperature scaling and ensemble-prompt entropy for calibration and uncertainty. We evaluate four medical VLMs on two cohorts (MARIO, OLIVES) across retrieval modes, neighbour counts, ablations, cross-cohort transfer and calibration — 238k logged predictions. Retrieval is decisive: without it the backends collapse onto one class (73% of decisions, κ≤0.18), whereas a single retrieved precedent raises κ by +0.26 on average (best leakage-free setting 0.712). Two controls show why: a caption query whose neighbours are label-matched by construction reaches κ=1.000, and the fitted temperature saturates its bound in every run while the confidence–accuracy gap widens from +0.385 to +0.569. Retrieved-label copying, not longitudinal reasoning, explains the headline numbers such pipelines report; we release the protocol with both controls as a baseline for longitudinal VLM evaluation.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/MedAgent_010.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: https://openreview.net/forum?id=5alIrUwYol
BibTex
@InProceedings{ZegRac_RetrievalAugmented_MICCAISAT2026,
author = { Zeghlache, Rachid AND Al Musleh, Mohamed AND Ahmed, Rehan AND Oroumchian, Farhad AND Alshryda, Sattar AND Elserafi, Ahmed AND Abdulaziz, Nidhal},
title = { { Retrieval-Augmented Longitudinal Reasoning for OCT Change Detection and Forecasting: A Trajectory-Aware Multimodal RAG Framework with Calibrated Vision-Language Models } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17263},
month = {pending},
page = {pending}
}
