Abstract

Parkinson’s disease exhibits heterogeneous motor presentations, yet data-driven subtyping is sensitive to modeling choices and often omits assignment uncertainty. We developed a posterior-aware Bayesian Gaussian mixture modeling framework using $n=29{,}366$ longitudinal MDS-UPDRS-III assessments from $1{,}847$ PPMI participants. A predefined search over 2{,}912 configurations selected a five-state representation, and patient-block bootstrap assessed component-occupancy stability. Model-conditioned posteriors assigned 99.5\% of assessments to a high-confidence category (\textsc{Textbook}) and 0.5\% to intermediate-confidence boundary assignments (\textsc{Chimera}). Cross-granularity correspondence between five- and eight-component solutions was strong ($V$ {=} $0.945$). In exploratory longitudinal analyses, bradykinesia scores most frequently improved prediction of subsequent axial scores. Motor-state assignments were associated with DaTSCAN striatal binding ratios ($n$ {=} $1{,}839$) and small-magnitude FreeSurfer subcortical volume differences ($n{=}1{,}706$; $\eta^2{=}0.005$–$0.010$; 13/25 ROIs FDR-significant). Applying the fixed scaler and BGMM to BioFIND ($n{=}310$) without refitting yielded 99.7\% high-confidence assignments. These findings support a reproducible visit-level motor-state representation with complementary imaging correlates. They do not establish five stable biological patient subtypes; repeated measures, possible severity confounding, and uncalibrated model posteriors limit interpretation. \keywords{Parkinson’s disease \and Motor state modeling \and Bayesian mixture models \and DaTSCAN imaging \and Structural MRI \and Uncertainty quantification \and Multi-scale clustering}

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/4053_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

https://github.com/CVC-Lab/posterior-aware-pd-phenotyping

Link to the Dataset(s)

PPMI: https://www.ppmi-info.org BioFIND / AMP-PD: https://www.amp-pd.org

BibTex

@InProceedings{TirHar_PosteriorAware_MICCAI2026,
        author = { Tirhekar, Harsh AND Yadav, Priyanshi AND Bajaj, Chandrajit},
        title = { { Posterior-Aware Motor Phenotyping with Multimodal Imaging Validation in Parkinson’s Disease } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16895},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    The main contribution of the paper is a posterior-aware motor phenotyping framework for Parkinson’s disease using Bayesian Gaussian Mixture Models (BGMMs) applied to large-scale MDS-UPDRS-III data (n=29,366 assessments) from the PPMI database. The approach explores 2,912 hyperparameter configurations with bootstrap validation to derive a robust k=5 phenotype solution, incorporating uncertainty quantification via posterior probabilities. The framework introduces a triage system (Textbook/Chimera/Ambiguous), evaluates temporal relationships via Granger causality, and validates phenotypes using DaTSCAN SPECT and structural MRI. External generalization is demonstrated on the BioFIND cohort without model refitting.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1) The hyperparameter sweep is impressive, along with bootstrap validation that improves statistical reliability. 2) Assigning posterior likelihoods as opposed to clustering subtypes is innovative and may be clinically useful. 3) Validating the approach against DATSCAN SPECT and FreeSurfer lends support to the biological plausibility of the approach. 4) External validation without refitting is a significant contribution.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1) One of the major weaknesses is the assumption of independence in samples, which is not true, as there are repeated measures per participant. This approach may inflate confidence. 2) BGMM, in and of itself, is not innovative. 3) Observed effect sizes are small, suggesting that the significance is driven by the number of samples rather than true biological separation. 4) The threshold chosen for posterior likelihoods is arbitrary with no clear justification. 5) Did the authors compare the model stability with random seeds or other clustering approaches? 6) How do the authors envision using the proposed model in practice, as the proposed model is computationally intensive and may not be practical?

  • Please rate the clarity and organization of this paper

    Poor

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not provide sufficient information for reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    No ground truth validation in BioFind, overinterpretation of Granger Causality, and arbitrary posterior thresholds dampen the enthusiasm of the proposed approach.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    Although I still doubt that the model can be used practically and also the question about validity remains, i think the model has enough points to justify publication to develop further.



Review #2

  • Please describe the contribution of the paper

    The authors propose a posterior-aware phenotyping framework that bridges data-driven motor clustering with multimodal neuroimaging analysis, using Bayesian Gaussian mixture models to differentiate subgroups with a large cohort of patients diagnosed with Parkinson’s disease. The methodology addresses 4 gaps in previous methods with validation in an independent dataset.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    A major strength of the paper is validation against both an independent Out-of-Cohort dataset and also with within cohort imaging data both of which add significant weight to their conclusions.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    No major limitations in the paper.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (5) Accept — should be accepted, independent of rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    Good out of cohort and multi-method validation.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The authors have adequately responded to my comments.



Review #3

  • Please describe the contribution of the paper

    This paper proposes a posterior-aware motor phenotyping framework for Parkinson’s disease based on Bayesian Gaussian mixture models (BGMMs) applied to large-scale longitudinal MDS-UPDRS-III assessments. The authors attempt to extend standard clustering studies by combining a large-scale hyperparameter sweep with bootstrap-based stability analysis, posterior-based uncertainty quantification that categorizes samples into Textbook, Chimera, and Ambiguous cases, multi-scale reconciliation between k=5 and k=8 solutions, temporal analysis using Granger predictability, and biological validation using DaTSCAN SPECT and structural MRI. The paper also reports external zero-shot generalization on BioFIND without refitting.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    1.The study is broad and technically ambitious.
    The paper combines clustering, posterior-based uncertainty analysis, hierarchical reconciliation, temporal analysis, multimodal imaging validation, and external evaluation, making it more elaborate than a standard single-result clustering study. 2.The configuration sweep is substantial.
    The search over 2,912 BGMM configurations, together with bootstrap validation, is a serious attempt to address the common sensitivity of clustering results to modeling choices. 3.The posterior-aware framing is conceptually interesting.
    Introducing Textbook, Chimera, and Ambiguous cases is more nuanced than hard assignment alone and is a potentially useful conceptual extension for phenotyping studies. 4.The use of DaTSCAN and MRI for downstream validation is a positive aspect.
    The attempt to relate the derived groups to imaging differences makes the work more translationally oriented than a purely statistical clustering paper.
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
    1.The central phenotype claim is built on a surprisingly narrow input representation.
    Despite the very broad presentation of the paper, the actual phenotype discovery step is based almost entirely on the 34 MDS-UPDRS-III motor items. MRI and DaTSCAN are not used to derive the clusters, but only for downstream validation. This means that the claimed richness of the phenotyping framework is much greater than the richness of the input information actually used to generate the subtypes. In my view, this is a fundamental limitation: the paper is performing a highly elaborate partitioning analysis on a relatively constrained and highly correlated motor rating space. 2.Extensive exploration over k does not solve the problem that the feature space itself is limited.
    The paper emphasizes the sweep over 2,912 BGMM configurations as a major strength. However, varying k, covariance type, and priors within a narrow motor-score space does not by itself make the result more biologically meaningful. At the end of the day, the model is repeatedly partitioning the same restricted clinical representation. This makes the analysis look richer than the underlying input actually is. 3.I am not convinced that the selected k=5 solution should be interpreted as five clinically meaningful phenotypes.
    The paper shows that k=5 is stable under the chosen internal clustering criteria and bootstrap protocol, but this is not the same as establishing that five is the most biologically or clinically grounded granularity. In fact, the paper itself notes that higher effective component counts occur frequently across the sweep, even if they score lower on the composite metric. The evidence therefore supports, at most, that k=5 is a convenient and stable partition under the authors’ scoring framework—not that Parkinson’s disease motor heterogeneity is naturally organized into five macro-phenotypes. 4.The reported groups remain readily interpretable as severity- or symptom-loading strata rather than true phenotypes.
    The resulting labels, such as Mild-Axial, Severe-Axial, Moderate-Tremor, and Severe-Tremor, strongly suggest a decomposition along severity and symptom burden axes. This makes it difficult to accept the stronger claim that the method has uncovered qualitatively distinct phenotypes. A much more convincing argument would be needed to rule out the simpler interpretation that the BGMM is discretizing a continuous motor severity manifold. 5.The visit-level design undermines the interpretation of the clusters as stable patient-level subtypes.
    The analysis uses 29,366 assessments from 1,847 patients, with repeated longitudinal visits per patient. Under this setup, the clustering may be capturing disease state transitions over time rather than stable phenotypic subgroups. The paper acknowledges the visit-level iid limitation, but this is not a peripheral issue—it directly challenges the meaning of the discovered clusters. 6.The uncertainty-aware contribution is weakened by an implausibly decisive empirical result. The paper emphasizes posterior-aware triage, yet 99.5% of samples are classified as Textbook, 0.5% as Chimera, and 0% as Ambiguous, with a mean posterior gap of 0.995.Given the known clinical heterogeneity of Parkinson’s disease, such an overwhelmingly decisive result is difficult to interpret as reassuring. It more naturally raises concern that the representation and modeling assumptions are producing overly sharp separation. The paper does not seriously interrogate this possibility. 7.Several downstream interpretations are overstated relative to what the analyses support. Granger predictability does not establish causal ordering or disease-mechanism hierarchy, yet parts of the discussion read as if it supports such interpretations. Likewise, the multi-scale hierarchy result is mathematically tidy, but the clinical value of this hierarchy is not demonstrated convincingly. These additional analyses make the paper appear more definitive than it actually is. 8.Overall, the paper is much stronger in analytical breadth than in conceptual validity.
    The study is elaborate, but the extra layers of analysis do not rescue the main weakness: the core subtype claim is not well established from the limited input representation and the current interpretation framework.
  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    The paper is clearly the result of substantial effort and contains several thoughtful technical components. However, I believe the current presentation substantially overstates what can be concluded from the analysis. In particular, the paper should more directly acknowledge that the subtype discovery is based only on motor assessment items, while the imaging modalities are used only for downstream validation. Under this design, a stable clustering solution should not automatically be treated as evidence for clinically meaningful phenotypes. I would strongly encourage the authors to more carefully distinguish between statistical partition stability and biological/clinical phenotype validity, to more directly address the possibility of severity-driven clustering, and to moderate several of the stronger clinical interpretations. Figure readability could also be improved, as several panels are visually dense.

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    The paper is clearly the result of substantial effort and contains several thoughtful technical components. However, I believe the current presentation substantially overstates what can be concluded from the analysis. In particular, the paper should more directly acknowledge that the subtype discovery is based only on motor assessment items, while the imaging modalities are used only for downstream validation. Under this design, a stable clustering solution should not automatically be treated as evidence for clinically meaningful phenotypes. I would strongly encourage the authors to more carefully distinguish between statistical partition stability and biological/clinical phenotype validity, to more directly address the possibility of severity-driven clustering, and to moderate several of the stronger clinical interpretations. In addition, several figures are visually dense, and some text/rendering appears degraded, which makes the presentation harder to follow. Improving figure readability and overall visual presentation would substantially strengthen the paper.

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (1) Strong Reject — must be rejected due to major flaws

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    I recommend Strong Reject. While the paper is ambitious and technically elaborate, I do not find its central claim convincing. The main issue is that the proposed phenotype discovery is based almost entirely on the 34 MDS-UPDRS-III motor items, which form a relatively narrow and highly correlated clinical feature space. The paper then performs an extensive clustering sweep over this restricted input space and interprets the selected k=5 solution as evidence for five meaningful Parkinson’s disease motor phenotypes. I do not think the evidence supports that conclusion. At most, the results show that k=5 is a stable partition under the authors’ chosen clustering metrics and bootstrap protocol, not that five is the correct or clinically grounded phenotype granularity. This concern is compounded by the fact that the resulting groups are still easily interpretable as severity- or symptom-loading strata, and by the visit-level design, which makes it difficult to distinguish stable patient-level phenotypes from longitudinal disease states. The uncertainty-aware framework is also much less convincing than advertised, given the almost perfectly decisive posterior assignments. Finally, the imaging results, while interesting, validate the partitions only after the fact and do not solve the core interpretational problem of what the clusters fundamentally represent. For these reasons, I view the paper as substantially over-interpreting a statistically stable clustering exercise, and I would score it well below the acceptance threshold in its current form.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Reject

  • [Post rebuttal] Please justify your final decision from above.

    The rebuttal is thoughtful and addresses my concerns directly, but it does not materially change my final assessment of the submitted paper. In particular, I remain unconvinced by the central subtype claim. The authors now clarify that the visit-level design is intended to characterize phenotypic states across the disease continuum, but this does not resolve the core interpretational issue that the paper still presents the selected clusters as clinically meaningful phenotypes. The distinction between stable patient-level subtypes and longitudinal disease states therefore remains insufficiently resolved.

    I also remain unpersuaded that the selected k=5 solution should be interpreted as evidence for five meaningful Parkinson’s disease motor phenotypes. The rebuttal argues that M1 and M3 have similar total severity but qualitatively different profiles, which is helpful, but it still does not fully overcome the broader concern that the clustering is derived from a narrow and highly correlated motor-score representation. In my view, the paper continues to support a statistically stable partition of motor presentations more strongly than a clinically grounded phenotype taxonomy.

    In addition, several rebuttal points rely on new or expanded analyses that are not part of the submitted manuscript, such as the proposed LME analysis and additional sensitivity analyses. While these points are useful context, they do not change the evidentiary strength of the paper in its current submitted form. The imaging validation remains interesting, but it is still downstream validation rather than evidence that the motor-only clustering input is sufficient to justify the stronger biological and clinical interpretation.

    Overall, I appreciate the authors’ effort and agree that the work is ambitious. However, I do not think the rebuttal sufficiently resolves the core concerns about limited input representation, the interpretation of k=5, the visit-level design, and the paper’s tendency toward overinterpretation. For these reasons, I keep my final decision as Reject.



Author Feedback

We thank the Reviewers and Area Chair (AC) for their constructive feedback. R1 highlights the “impressive hyperparameter sweep” and innovative “posterior likelihoods.” R2 recognizes the “strong out-of-cohort validation.” R3 notes the study is “technically ambitious” and “conceptually interesting.” We address the major concerns below.

1.Visit-level Independence vs. Patient-level Subtypes [R1, R3, MR] The visit-level design serves as a characterization of Phenotypic States across the disease continuum. Our goal is to model the distribution of clinical presentations at any assessment. However, our data demonstrates high patient-level stability: for patients with >=5 visits (n=2,014), the mean cluster stability is 84.6%, and stability within the Tremor vs. Axial families (among symptomatic visits) is 89.7%. This proves that discovered clusters map to stable phenotypic trajectories, not just transient states. To address the independence assumption in p-values, we will include Linear Mixed-Effects (LME) models in the camera-ready version to confirm imaging differences remain significant after accounting for patient-level random effects.

2.Phenotypes vs. Severity Strata [R3, MR]: While severity is a component, our results demonstrate qualitative profile divergence independent of burden. M1 (Sev-Tremor) and M3 (Sev-Axial) occupy similar total MDS-UPDRS-III severity ranges but exhibit fundamentally different motor signatures (Fig 2a). Crucially, our novel DaTSCAN Asymmetry Index (AUC=0.901) distinguishes these groups with high precision, proving that motor-defined phenotypes map to distinct, lateralized patterns of neuro-degeneration, not just global severity levels.

3.Narrow Input Representation (Motor-Only Clustering) [R3, MR]: Discovery on motor scores ensures clinical interpretability and actionability. The richness of the framework comes from proving these clinically-derived groups have independent neurobiological correlates in DaTSCAN (p<10^-8) and MRI (p<0.01 FDR). This cross-modal validation (noted as a strength by R1, R2) shows motor presentation is a robust proxy for pathology. Furthermore, these posterior-aware labels can serve as “soft-targets” for future deep learning models to learn imaging-based signatures of PD progression, directly contributing to the MICCAI imaging community.

4.Decisiveness of Posteriors (99.5% Textbook) [R3, R1]: The high decisiveness reflects the robust domain structure of MDS-UPDRS-III and BGMM convergence on a massive dataset (n>29k). This is consistent with the well-known “Tremor vs. PIGD” separation in clinical literature. Sensitivity analysis shows that varying the triage threshold from 0.7 to 0.9 maintains the same clinical destinations and lateralization gradients, confirming the robustness of the discovered structure.

5.Effect Sizes and p-values [R1]: Small effect sizes (η²) are expected in de novo PD where changes are subtle. The significance lies in the multi-modal convergence (SPECT + MRI + BioFIND) which collectively validates phenotypic boundaries across three independent data sources.

6.Clarity and Computational Intensity [R1, R3]: We will simplify Figure 2 and provide a detailed sweep pseudocode. While discovery is computationally intensive, inference for new patients is near-instantaneous, making the model practical for clinical trial enrichment and bedside triage.

Summary: Our framework bridges clinical assessment with neurobiology. By addressing the “state vs. subtype” distinction and emphasizing qualitative imaging divergence (AUC=0.901), we provide a robust path for acceptance in practice.




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    Strong point is the validation approach, however the majority of reviewers recommend reject. Please addres the concerns, especially the limited input data and feature space and the risk over overinterpretation (e.g. the interpretation of k=5, the use of visit as independent samples). The conclusions may be an overstatement of the experiments, the authors may want to change this.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The rebuttal is thoughtful and address most reviewers’ concerns. Not all reviewers are convinced by the practical use and on the central claim, but the majority of the reviewers recommend acceptance.



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    This paper presents a posterior-aware phenotyping framework for Parkinson’s disease using Bayesian mixture modelling, uncertainty-aware phenotype assignment, and multimodal imaging validation. Reviewers highlighted several strengths, including the extensive hyperparameter exploration, uncertainty-aware phenotyping strategy, biological validation using DaTSCAN and MRI, and external validation on an independent cohort.

    The main concerns focused on the interpretation of discovered phenotypes, limited feature space used for clustering, and risk of over-interpreting statistically stable clusters as biologically distinct phenotypes. The rebuttal clarified the intended focus on phenotypic states across the disease continuum, provided additional evidence for longitudinal stability, and better positioned the imaging analyses as biological support for the discovered structure. While some concerns remain regarding phenotype interpretation and distinction between disease states and stable subtypes, the paper offers a comprehensive and well-validated framework with strong external validation and clinically relevant imaging analyses.



Meta-review #3

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Reject

  • Please justify your recommendation.

    This paper has very diverse scores, ranging from strong rejection to acceptance. After reading the rebuttal and the reviewer’s comments, I agree with the negative reviewer that the main claim may be misleading. One of the reviewers, indicating acceptance after rebuttal, still has a concern about the model’s usability in a real clinical setting. As noted by the reviewer, the central concept cannot be changed by rebuttal only. So I am leaning toward rejecting the paper.



back to top