List of Papers Browse by Subject Areas Author List
Abstract
Longitudinal neuroimaging is essential for studying brain aging and disease progression. Most existing approaches track raw change over time, without reference to normative population trajectories. Cross-sectional normative models offer a principled alternative by contextualising an individual’s brain against healthy population norms, but their application to longitudinal inference raises an underexplored question: how reliably can individual deviations from normative trajectories be distinguished from stochastic measurement error? We address this using the Longitudinal Deviation Index (LDI), which quantifies reliability as the ratio of true inter-individual residual slope variance to variance expected purely from measurement error. Leveraging a normative model pre-trained on over 6,800 healthy individuals, we derived theoretical LDI estimates from test–retest data and validated them against empirical longitudinal data from cognitively normal adults. Within an 8-year follow-up window (LDI < 1), individual deviations are indistinguishable from measurement noise, confounding valid longitudinal inference. Beyond 8 years (LDI > 1), cumulative longitudinal variations systematically surpass the noise floor, reflecting genuine biological progression. These results clarify the limits of longitudinal interpretation using cross-sectional normative models and provide a principled framework for noise-aware inference.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/3319_paper.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to the Code Repository
N/A
Link to the Dataset(s)
N/A
BibTex
@InProceedings{JiWen_Longitudinal_MICCAI2026,
author = { Ji, Wenxing AND Long, Jianqiao AND Guan, Qiwen AND Li, Xinyu AND Jiang, Shouyong AND Li, Jichun AND Wang, Yujiang},
title = { { Longitudinal Reliability of Cross-Sectional Normative Models Under Measurement Error } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 16887},
month = {September},
page = {pending}
}
Reviews
Review #1
- Please describe the contribution of the paper
This paper studies the reliability of longitudinal inferences derived from cross-sectional normative models in neuroimaging. The authors introduce the Longitudinal Deviation Index (LDI), defined as the ratio between inter-individual variance in residual slopes and the variance expected from measurement error, to quantify whether observed longitudinal changes reflect true biological variation or noise. Using a pre-trained normative model and data from cognitively normal subjects in the OASIS-3 dataset, the authors derive a theoretical formulation of LDI and validate it empirically across varying follow-up durations and sampling densities. The results show that, for follow-up periods shorter than approximately eight years, longitudinal variability is largely dominated by measurement error, while longer follow-ups reveal deviations not captured by cross-sectional models. The work provides a quantitative framework for assessing the validity of longitudinal interpretations based on normative modeling.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The paper addresses a relevant and timely problem in medical image analysis, namely the misuse or over-interpretation of longitudinal signals derived from cross-sectional normative models. The proposed LDI metric is well-motivated, interpretable, and grounded in statistical principles, offering a practical way to disentangle biological signal from measurement noise. A notable strength is the combination of theoretical derivation and empirical validation, with good agreement between predicted and observed behavior. The use of test–retest data to estimate measurement error is appropriate and strengthens the methodological rigor. The findings are insightful and somewhat counterintuitive, particularly the observation that increasing the number of time points does not compensate for short follow-up duration. The analysis of regional heterogeneity further enhances the relevance of the work for neuroimaging applications.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.In Figure 1a, it would be recommended to use non-parametric fit, such as loess fit, to show if the association between measurement error and age is linear or not.
2.In Figure 1a, it would be nice to show somewhere the measurement error in absolute value scale, without taking the log.
3.The assumption of stable normative centiles over time is strong and may not hold even in healthy populations, potentially confounding interpretation of residual changes.
4.The interpretation of high LDI values as evidence of model inadequacy is somewhat overstated, as such effects could also arise from biological heterogeneity or violations of modeling assumptions.
5.Empirical validation is limited to a subset of sampling scenarios (3 and 5 time points), leaving other configurations untested.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The proposed Longitudinal Deviation Index (LDI) is a clear and principled metric with intuitive interpretation, and the combination of theoretical derivation with empirical validation strengthens the credibility of the approach. The finding that follow-up duration is more critical than sampling frequency is particularly insightful and has practical implications for study design in longitudinal imaging.
However, the framework relies on strong assumptions that most notably linear residual trajectories and stable normative centiles. This is not fully justified and may not hold in real-world settings. These assumptions directly impact the interpretation of LDI and raise concerns about whether high LDI values reflect model limitations, biological variability, or assumption violations.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #2
- Please describe the contribution of the paper
The paper introduces a quantitative metric to capture longitudinal individual variation in population models.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
It aims to capture longitudinal variation and doesn’t overstate potential use-cases of this approach.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
- is it safe to assume residual trajectories are linear?
- Maybe i got lost in notice but how do we know/quantify the true measurement error? Or at least how can technical noise be dissociated from true biological noise over time?
- Is 3 observations really enough to estimate the true biological variance and is it approrpriate to assume that this is linear?
- the statement that measurement error is age-indepedent should probably come with a very clear warning label that this is only the case in the age range investigated and over the limted time-frame of the test-retest data
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
I don’t think its a particularly exciting contribution and for a full citable paper I would very much like to see more robustness in different datasets and age-ranges, but the paper does not entirely overstate its contribution.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #3
- Please describe the contribution of the paper
This paper studies the relationship between measurement error at single time points and longitudinal error in characterizing brain change over time. Empirical assessments based on the theoretical model proposed in the paper suggests that longitudinal changes over a period of up to 8 years can be reliably characterized from cross-sectional normative modelling.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The major strengths of the paper are the theoretical modelling for longitudinal change over time, and the careful experiments that were carried out to assess the impact of measurement error on image-derived phenotypes.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The paper is without major weaknesses, and could be adopted and extended in several valuable directions.
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission has provided an anonymized link to the source code, dataset, or any other dependencies.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(5) Accept — should be accepted, independent of rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The paper provides an interesting theoretical model and a practical evaluation and validation of the model that is broadly applicable to many applications of medical image computing.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #4
- Please describe the contribution of the paper
- Longitudinal Deviation Index (LDI) as a principled metric for quantifying the reliability of longitudinal deviations derived from cross-sectional normative models is introduced
- both theoretical derivations and empirical validation shows that short-term residual slope variability is largely attributable to measurement error
- the work clarifies the conditions under which cross-sectional normative models can support longitudinal inference and highlights their limitations over extended follow-up periods
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
- by evaluating the reliability of longitudinal inference from widely used cross-sectional normative models it addresses an important and timely problem in longitudinal neuroimaging
- theoretical analysis is combined with empirical validation, providing mutually supportive evidence for the proposed Longitudinal Deviation Index (LDI)
- proposes a clear and practically relevant indicator (LDI) for distinguishing true longitudinal signal from measurement noise
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
- empirical validation is limited to a single-site cohort of cognitively normal subjects, which restricts generalizability to multi-site studies and, most critically, clinical populations
- several key assumptions, particularly normally distributed slopes and random measurement error, may be difficult to justify in real longitudinal neuroimaging settings with systematic scanner-related effects
- writing needs substantial improvements; i.e. important concepts and terminology are imprecisely defined or used inconsistently, reducing interpretability (see specific comments)
- practical utility for detecting abnormal trajectories or evaluating specific imaging-derived phenotypes is not yet fully demonstrated
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not provide sufficient information for reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
1.The manuscript states: “We assume that population slopes (b_i) follow a normal distribution, (b_i \sim N(\epsilon, \vartheta_b^2)). Similarly, the measurement error (\omega_{ij}) is normally distributed with mean 0 and variance (\vartheta^2).” These distributional assumptions appear to have been assessed using a single-site dataset. However, the manuscript does not sufficiently describe how these assumptions were empirically verified. A more rigorous justification is needed, including explicit normality diagnostics, goodness-of-fit assessments, and discussion of whether such assumptions are expected to generalize beyond the specific single-site cohort analyzed.
2.The statement that “applying linear regression to individuals with at least three observations allows us to estimate the true biological variance while isolating it from measurement error” relies critically on the assumption that measurement error is random and independent. This assumption may be problematic in longitudinal neuroimaging settings, where systematic sources of variation—such as scanner drift, hardware upgrades, site effects, acquisition protocol changes, and preprocessing inconsistencies—can induce structured rather than random error. The manuscript should more carefully discuss these potential violations and their implications for the validity of variance decomposition.
3.The manuscript states that “a low LDI indicates that the residuals … that normative model has accurately encapsulated the true longitudinal progression.” The term “true progression” is insufficiently defined and may be conceptually ambiguous. It is unclear whether this refers to biological aging trajectories, disease-related change, or adherence to the expected normative centile trajectory. A more precise formulation would improve interpretability; for example, stating that a low LDI suggests the individual’s centile remains consistent with the trajectory predicted by the normative model.
4.Figure 2 is not explicitly referenced in the main text, which makes its role in the narrative unclear. In addition, the caption does not adequately specify whether the figure presents theoretical results, empirical validation results, or both. The figure should be clearly introduced in the text and its caption revised to state the source and interpretation of the displayed results.
5.The phrase “complex inter-subject variability” is overly broad and lacks scientific specificity. The authors should clarify what dimensions of variability are intended (e.g., heterogeneity in baseline values, slope differences, nonlinear trajectories, demographic effects, or measurement-related variability).
6.A strength of the manuscript is that it presents theoretical and empirical evidence side-by-side, yielding broadly consistent conclusions regarding the LDI and thereby indirectly supporting some of the underlying assumptions. However, the empirical validation is not fully systematic. As acknowledged by the authors, “First, limited empirical data prevented the validation of scenarios involving 7 and 9 observations.” Moreover, validation was restricted to cognitively normal subjects, whereas the principal clinical relevance of the framework likely lies in populations deviating from normative aging trajectories. Extending the analysis to pathological or at-risk cohorts would substantially strengthen the manuscript and better demonstrate the practical utility of the LDI, particularly for evaluating the suitability of specific imaging-derived phenotypes (IDPs).
7.The authors appropriately acknowledge that the study was limited to single-site data. This is an important limitation because the assumptions underlying LDI construction may be more readily violated in multi-site settings due to “scanner drift, hardware upgrades, site effects, acquisition protocol changes, and preprocessing inconsistencies.” Since many contemporary neuroimaging studies are multi-center, the manuscript would benefit from a more detailed discussion of how “robust” the proposed framework is to such sources of heterogeneity and whether harmonization strategies would be required.
8.Several key terms are used inconsistently throughout the manuscript, which reduces conceptual clarity. Examples include: “the cross-sectional model”, “normative model”, and “cross-sectional normative model”; as well as “the true longitudinal progression”, “longitudinal structure”, “true longitudinal dynamics”, and “longitudinal biological variation.” The terminology should be standardized, and each concept clearly defined, to ensure consistency and avoid ambiguity.
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The manuscript tackles a relevant challenge in longitudinal neuroimaging and proposes a potentially useful reliability metric (LDI), and the combination of theoretical and empirical analyses is a clear strength. However, empirical results are restricted to a single-site cohort of cognitively normal subjects, reducing generalizability. Given the vast amount of publicly available longitudinal neuroimaging data, such analyses are readily possible. Further, key assumptions on random measurement error and distributional forms require stronger justification. Finally, writing needs substantial improvements. While promising, broader validation and clearer methodological support and presentation are needed before acceptance.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Author Feedback
We sincerely thank the Meta-reviewer and all reviewers for their insightful and constructive feedback. We are honored by the “Early Accept” decision and encouraged by the reviewers’ recognition of our work’s merit, importance, and methodological rigor. In the camera-ready version, we will fully integrate the suggested improvements.
1.Addressing Key Assumptions (Linearity and Normality)
Several reviewers questioned the assumption of linear residual trajectories and the normal distribution of slopes. We acknowledge that biological brain changes can be non-linear. However, for this foundational study, we utilized a linear first-order approximation to provide a simplified, interpretable, and tractable analytical framework for quantifying longitudinal reliability (LDI). We also confirm that normality diagnostics were conducted for measurement error and slopes, confirming our assumptions; these details were omitted in the initial submission due to page constraints, but will be clarified in the revision.
2.Multi-site Generalizability and Structured Noise
Reviewer #4 raised concerns regarding single-site limitations and systematic scanner effects. We clarify that the OASIS-3 dataset used in this study encompasses longitudinal data collected across multiple scanners. We apologize for the insufficient description in the original text and will provide a comprehensive site description in the final version. Furthermore, to mitigate systematic biases like scanner drift, we utilized site-specific parameters during normative modeling to calculate calibrated IDPs, effectively decoupling measurement error from cross-sectional hardware differences. We will discuss harmonization strategies for broader multi-center applications in the Discussion Section.
3.Measurement Error and Noise Floor
Regarding the dissociation of technical noise from biological signal, we utilized 15-day test-retest data to quantify measurement error, assuming biological stability over such a short interval. This allows us to establish a “noise floor.” An LDI > 1 signifies that observed deviations surpass this noise floor, reflecting genuine biological progression not captured by the cross-sectional model. We will explicitly discuss the limitations of this assumption in the Discussion.
4.Clinical Utility and Future Directions
While our current empirical validation was restricted to cognitively normal subjects to establish a rigorous noise-propagation baseline, we agree that LDI’s primary utility lies in identifying abnormal trajectories in pathological cohorts. In our future work, we will carefully consider the practical utility of LDI for evaluating biomarkers in clinical settings.
5.Clarification and Presentation Improvements
We will implement the following specific updates: Visualizations: Update Figure 1a with a non-parametric loess fit and include absolute physical units on the axis labels to enhance interpretability. Citations: Explicitly reference Figure 2 in the text and revise its caption to distinguish between theoretical and empirical results. Terminology: Conduct a linguistic audit to standardize terminology and replace broad phrases with specific biological descriptions.
We believe these revisions will significantly strengthen the manuscript’s clarity and impact. We look forward to presenting our refined work at MICCAI 2026.
Meta-Review
Meta-review #1
- Your recommendation
Provisional Accept
- Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.
Reviewers summarized strengths of your work, novel aspects and noted that your solution may be of potential interest. They all agree on the merit and importance of this research, which is the evaluation of reliability and validity of longitudinal modeling given cross-sectional data (statistics on mixed-effects modeling shows that results can be quite different if not using true longitudinal input data). Reviewers also mentioned the analysis of regional heterogeneity which seems of relevance to neuroscience, the insightful findings and the methodological rigor. of the analysis. Reviews also brought up some limitations such as limited empirical validation, the question of linearity of residual trajectories, and expected to see analysis of robustness on different datasets. Would the paper be accepted for MICCAI, authors are strongly encouraged to provide a revision that integrates suggestions for improvements that consider specific weaknesses as discussed in the reviews, in particular addressing details of limitations as brought up by reviewer #4.
