Abstract
Longitudinal management of multiple myeloma (MM) requires evaluating complete treatment trajectories rather than isolated clinical encounters. Existing prognostic models focus primarily on risk prediction and do not explicitly assess cumulative utility of sequential treatment decisions. In this study, we present TRAIL (Trajectory-Aware Reward Attribution and Inference for Longitudinal Multiple Myeloma Treatment Trajectories), a preference-based offline reinforcement learning framework that learns stepwise reward functions and estimates trajectory returns from longitudinal electronic health records (EHRs). We formulate MM management as a discrete Markov decision process and represent irregular visit intervals using Fourier temporal features and a causal transformer. A Bradley-Terry objective is then used to learn reward functions from outcome-derived trajectory preferences. The learned reward function is subsequently used for offline policy extraction via Fitted Q-Iteration (FQI) or Conservative Q-Learning (CQL). Experiments on a cohort of 575 MM patients demonstrate that TRAIL outperforms representative offline policy learning baselines, achieving a C-index of 0.766 and an AUC of 0.738. Clinical validation further shows that the predicted trajectory return remains associated with survival outcomes across multiple evaluation horizons and risk stratification.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/CaPTion_015.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: https://openreview.net/profile?id=~Tao_Chen33
BibTex
@InProceedings{CheTao_TRAIL_MICCAISAT2026,
author = { Chen, Tao AND Zhou, Chuan AND Wang, Yifan AND Hadjiiski, Lubomir AND Wilms, Matthias},
title = { { TRAIL: Trajectory-Aware Reward Attribution and Inference for Longitudinal Multiple Myeloma Treatment Trajectories } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17278},
month = {pending},
page = {pending}
}
