<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" ><generator uri="https://jekyllrb.com/" version="3.9.0">Jekyll</generator><link href="https://papers.miccai.org/miccai-2026/feed.xml" rel="self" type="application/atom+xml" /><link href="https://papers.miccai.org/miccai-2026/" rel="alternate" type="text/html" /><updated>2026-09-21T21:34:43-04:00</updated><id>https://papers.miccai.org/miccai-2026/feed.xml</id><title type="html">MICCAI 2026 - Open Access</title><subtitle></subtitle><entry><title type="html">3D Cerebrovascular Shape Completion from Biplane Angiography and CTA Prior</title><link href="https://papers.miccai.org/miccai-2026/0001-Paper2705" rel="alternate" type="text/html" title="3D Cerebrovascular Shape Completion from Biplane Angiography and CTA Prior" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0001-Paper2705</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0001-Paper2705">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;We propose a novel approach to bridge the resolution and dimensionality gap between CTA, which provides 3D vascular geometry but often misses small vessels, and biplanar 2D DSA, which offers higher spatial resolution but lacks 3D structural information, by formulating the problem as a 3D shape completion task.
Our method represents vasculature using a set of 3D Gaussians initialized from CTA-derived vessel geometry and augmented with additional spatial and opacity primitives seeded from two DSA projections. These Gaussians jointly encode geometry and attenuation and are optimized to fit the observed DSA images while remaining consistent with the original CTA anatomy.
Experiments on synthetic and clinical cerebrovascular data demonstrate improved 3D reconstruction of small vessel branches at submillimetric resolution, validating the use of DSA to complement missing CTA vessel anatomy.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/2705_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/2705_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: &lt;a href=&quot;https://papers.miccai.org/miccai-2026/supp/2705_supp.zip&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/supp/2705_supp.zip&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/janik-j/cta-dsa-fusion&quot;&gt;https://github.com/janik-j/cta-dsa-fusion&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;TopBrain CTA Dataset: &lt;a href=&quot;https://topbrain2025.grand-challenge.org/data/&quot;&gt;https://topbrain2025.grand-challenge.org/data/&lt;/a&gt;
Clinical CTA/DSA dataset (Brigham and Women’s Hospital): Not publicly available due to PHI constraints
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{JehJan_3D_MICCAI2026,
        author = { Jehkul, Janik AND Frisken, Sarah AND Gopalakrishnan, Vivek AND Rueckert, Daniel AND Haouchine, Nazim},
        title = { { 3D Cerebrovascular Shape Completion from Biplane Angiography and CTA Prior } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16889},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This work presents a Gaussian splatting-based method to reconstruct 3D mesh from bi-plane DSA, using the pre-operative CTA as the prior.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The proposed method seems to outperform the SOTA methods based on qualitative and quantitative assessments. 
2.Loss functions are tailored to the task.
3.The ablation studies are comprehensive and clearly demonstrate the contribution of each component.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Only a subset of TopBrain CTA data was used, but the authors didn’t provide the reason why they exclude the rest. 
2.The DSA data were synthesised from CTA, which could be unreliable. In reality, CTA can visualise full cerebral vasculature whereas DSA only visualise part of it, because of the difference in the contrast injection methods (intravenous vs intra-arterial). In addition, there might be variation due to contrast agent dispersion. The authors should provide further information on how the DSA data synthesis was performed. 
3.The method was tested on only three patients with acquired DSA and CTA. The projection of the reconstructed mesh seemingly agrees well with the original DSA in Patient 1.However, visual assessment based on the provided 3D meshes is very difficult for Patient 2 &amp;amp; 3 and unable to confirm successful reconstruction.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors claimed to release the source code and/or dataset upon acceptance of the submission.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The proposed method relies on CTA/DSA registration for real world application. How reliable is the registration used in this work and how significant would registration errors affect the final result? 
2.The method employs multiple loss terms—could the authors elaborate on how their weights were balanced? Table 2 shows that combining all the designated losses does not achieve the best performance. Please provide further discussion.
3.(Figure 4) Why is the internal carotid artery missing in the incomplete mesh in every patient? As the main artery, it should be easily identified at initialisation.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;I am not convinced that the proposed method has achieved the claimed sucess.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors have addressed most of the previous concerns. Importantly, they have agreed to update Fig 4 to improve visualisation and to support the claims. While the small in vivo dataset of three patients makes the validation qualitative and speculative, it might be sufficient for proof of concept. 
Overall, the recommendation is acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper proposes a 3D cerebrovascular shape-completion method that combines an incomplete CTA-derived vessel prior with two biplane DSA projections. The core idea is to represent the vasculature using 3D Gaussians initialized from both the CTA mesh and DSA-derived residual context points, and then optimize them with a rendering loss, a false-negative vessel loss, and a prior-anchoring regularizer.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;The paper addresses a clinically meaningful problem. CTA provides 3D structure but, as the authors stress, often misses small distal vessels, while biplane DSA has higher spatial resolution but only 2D information. Thus, framing this as shape completion from a patient-specific CTA prior plus biplane DSA is a strong and well-motivated formulation&lt;/li&gt;
        &lt;li&gt;The method itself is technically interesting. The combination of dual-source Gaussian initialization, differentiable X-ray rasterization, a false-negative loss targeted at weakly supervised thin branches, and a prior-anchoring term is coherent and well aligned with the imaging physics and the stated goal of vessel completion rather than generic reconstruction. Overall, it feels like a quite interesting and novel approach&lt;/li&gt;
        &lt;li&gt;Quantitative results are strong on the synthetic benchmark. The method clearly outperforms the reported baselines in both 2-view and 3-view settings&lt;/li&gt;
        &lt;li&gt;The paper also includes meaningful ablation studies on losses, initialization strategy, and prior completeness&lt;/li&gt;
        &lt;li&gt;The paper includes an initial clinical demonstration. Although a bit limited, the three-patient qualitative study supports the claim that the method can recover distal branches visible in DSA but absent from CTA&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;The main weakness is the limited real-world validation. Quantitative experiments are performed on only 10 synthetic CTA scans with simulated incomplete priors, while the clinical study covers only 3 patients and is qualitative only. There is no paired clinical ground truth or quantitative evaluation on real patient data, so it is still unclear how robust the method is under realistic acquisition artifacts, contrast inhomogeneity, and registration error&lt;/li&gt;
        &lt;li&gt;A second weakness is that the novelty is somewhat incremental at the problem level, even if the specific combination is novel. Prior work has already studied two-view cerebral vascular reconstruction and learned priors for this setting (e.g., “Two Projections Suffice for Cerebral Vascular Reconstruction” by Frisken et al.). The novelty of the present paper lies more in the CTA-guided Gaussian splatting completion formulation than in introducing the two-view reconstruction problem itself&lt;/li&gt;
        &lt;li&gt;A more minor weakness is that reproducibility is only partial in the current version. The paper gives the overall formulation and some hyperparameters, but several implementation details appear under-specified in the main text, such as some initialization and threshold choices for the auxiliary terms and prior neighborhood construction&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors claimed to release the source code and/or dataset upon acceptance of the submission.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;I think it’s an interesting paper overall and well-written. I liked the idea of using a CTA prior to somewhat turn biplane DSA reconstruction into a completion problem rather than a purely unconstrained inverse problem. The ablations are helpful and the improvements on the synthetic benchmark are convincing.&lt;/p&gt;

      &lt;p&gt;The main issue for me is validation. The paper would be significantly stronger with either quantitative evaluation on clinical data or a more realistic simulation study that better captures registration error, residual subtraction artifacts, and contrast inhomogeneity.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;I am leaning weak accept because the paper presents a solid and clinically relevant methodological contribution: a well-motivated CTA-guided Gaussian-splatting formulation for completing cerebrovascular anatomy from only two DSA views. The quantitative gains over the reported baselines are substantial, and the ablations help support the design choices.&lt;/p&gt;

      &lt;p&gt;At the same time, the evidence for real-world clinical utility is still limited. The quantitative study is small and synthetic, and the clinical evaluation is qualitative on only three patients.&lt;/p&gt;

      &lt;p&gt;So to summarize, I find the idea strong and potentially impactful but I think the paper still sits near the threshold because the empirical validation does not yet fully match the ambition of the claims.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;I believe the authors have adequately addressed the concerns raised in my review. I am satisfied with the authors’ response.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Authors are proposing taking advantage of imaging methods with higher resolution to complement imaging methods with a lower resolution. In this case, completing the vasculature observed in CTA which in most hospital setups can go up to 0.5mm (or 0.83mm in the particular clinical evaluation of this paper), with the one observed in DSA acquisitions, the gold standard in many neurovascular diseases like stroke, which tend to go up to 0.2mm. Particularly helpful in treatment of small aneurysms, stroke and potentially calcification in distal vessels. Authors method to complete the vasculature uses dual initialization, one using a prior mesh from the CTA and DSA biplanar projections (improved with a third oblique view). Loss function includes a rendering component, a false negative vessel loss and a prior anchoring regularizer, to cover all the aspects of the method. Method is evaluated on both synthetic and clinical datasets (n=10 and n=3 respectively).&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Complete reconstruction of distal vessels having just 3 or even two additional DSA planes and extend the baseline CTA is a great piece of work. Not only because of the extension of diagnostic tools, but also because particular neurovascular diseases whose gold standard diagnostic image is DSA will benefit from having complete distal information for treatment planning.&lt;/p&gt;

      &lt;p&gt;The loss function covers important aspects of the proposed method: how good is the render, regularize on the priors and check false negative on distal branches. The term is well thought and parameters of the training objective are well defined (potentially tuned). False loss term makes sense since the original name and main method is the study on the contribution of sparse (biplanar) DSA, however is worth asking on the review side, is it the same parameter value when using the oblique view? Given the fact of how much it improves the results. Are the terms performing the same on the both biplanar and triplanar setups?&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;While the community broadly accepts demonstrations of clinical feasibility on small cohorts, the clinical evaluation (n=3) remains exclusively qualitative. Given that the paper explicitly claims clinical applicability and demonstrates recovery of 0.216mm vessels within 0.833mm CTA resolution, it would be interesting to disclose quantitative metrics for the clinical cases. Was the decision to keep this qualitative a consequence of annotation difficulty, or was there a deliberate methodological reason? Clarifying this would help readers assess the generalizability of the results.
The method’s performance somehow relies on accurate 2D-3D registration between DSA planes and the CTA-derived mesh. However, no registration error analysis or sensitivity study is provided. Understanding there are space limitations and authors covered quite well the complete explanation, a small discussion on how robust the pipeline is to misalignment would be good.
A small extension of the previous point is that you had space enough to cover a context point initialization ablation, however information of this table is limited. It is interesting why having less context points (only M1) perform better than uniform initialization.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors claimed to release the source code and/or dataset upon acceptance of the submission.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper presents a technically well explained method and a clinically motivated contribution at the using CTA and DSA imaging. The synthetic evaluation is well thought, the loss design is well motivated, and the gains from adding a third oblique plane are really notable. The contribution is novel and the clinical framing is, at least, relevant to diseases like stroke.
The score is Weak Accept given the absence of quantitative metrics (CD, HD) in the clinical evaluation needs justification, if it was a deliberate design decision or an annotation constraint should be explicitly stated. Registration quality between DSA can be commented/expanded, as well as how brief was table 3 commented.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The rebuttal addresses my main concerns directly. The justification for the qualitative-only clinical evaluation, that quantitative ground truth would require CTA acquired at near-DSA resolution, which is not standard in clinical practice, is a methodologically honest answer, and the reframing of the 10 TopBrain cases as the quantitative cohort and the 3 clinical cases as feasibility-under-real-DSA conditions is reasonable, but still quite small, worth noting. The registration response (XVR with sub-mm accuracy, combined with Eq. 5’s KNN-anchored prior absorbing sub-voxel imprecision) explains the design feature that confers robustness, but not explain at all how good does registration need to be, a full sensitivity study would not hurt. The loss-weight clarification at biplane vs triplane answered my question about whether the loss terms perform equivalently across setups The contribution looks good and clinically motivated.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;

  &lt;p&gt;We are glad reviewers found the problem “clinically meaningful” (R2); our method “novel”, “technically interesting” (R2), and “well explained” (R3) with a “strong and well-motivated” formulation (R2) and losses “tailored to the task” (R1). Reviews also found the results to “outperform[s] the SOTA methods” (R1) and the ablations to be “comprehensive” (R1) and that we overall presented a “great piece of work” (R3). 
We thank the reviewers for their comments and address the main ones below.&lt;/p&gt;

  &lt;p&gt;CLINICAL VALIDATION (R1,R2,R3): 
We presented a quantitative evaluation on 10 TopBrain cases using real CTA and synthetic DSA, including comparisons with SOTA methods, an ablation study, and a sensitivity analysis, along with qualitative results on 3 clinical cases. These additional clinical cases are provided to demonstrate feasibility under real clinical DSA acquisition conditions, not to serve as a statistically powered evaluation cohort. 
Moreover, quantitative clinical evaluation would require CTA acquisition at a resolution comparable to DSA (to capture tiny vessels), which is not widely acquired in clinical practice.
In addition, since our approach is a patient-specific optimization, all 10 cases contribute independently to the results in Tables 1-3, covering different anatomies and vascular variability.
Altogether, while we agree that a larger clinical evaluation is beneficial, we believe that our evaluation supports the claims of the paper and is overall on par with related methodological papers in this area.&lt;/p&gt;

  &lt;p&gt;SYNTHETIC DSA (R1): 
We agree CTA covers more vasculature than DSA. Our aim is not to match coverage but to use CTA as a prior to reach the sub-mm resolution that only DSA resolves, as stated in the introduction. We forward-project each CTA volume with TIGRE under a simulated cone-beam geometry and subtract mask from fill projections (Sec. 5.1). The clinical input I_v is a temporal MIP (Sec. 2.1), so both inputs reduce to a single static projection per view. R1 rightly notes that real DSA only depicts the injected territory due to IA vs IV contrast administration. The synthesis does not model this, but the clinical pipeline crops M_P to points projecting into DSA vessel masks across all views (Sec. 5.1.2), restricting the prior to the injected territory at input. We will clarify this in the final paper.&lt;/p&gt;

  &lt;p&gt;TOPBRAIN SUBSET (R1, R2): 
We had to use a subset of TopBrain to fine-tune DenoiseNet, leaving 10 cases for evaluation. No method sees test data at training time. We will state this split explicitly in the paper.&lt;/p&gt;

  &lt;p&gt;LOSS WEIGHTS (R1, R3): 
Each loss targets one error mode by design, so the full configuration is not best on every per-region metric. L_FN improves the missing region HD_ROI by 18%, L_PA the prior region CD_prior by 8%, combined best overall CD. λ_FN and λ_PA are identical at 2 and 3 views with effects more pronounced at 3 views (R3). λ_ssim differs (0 vs 0.07), as we empirically found that pure L1 works better at 2 views.&lt;/p&gt;

  &lt;p&gt;REGISTRATION (R1, R3): 
CTA to DSA alignment uses XVR [7] (sub-mm accuracy). The CTA prior is not fixed: Eq. 5 anchors centers to the local mean of their KNN prior points, and the squared norm penalizes large displacements, so sub-voxel imprecision is absorbed without prior drift.&lt;/p&gt;

  &lt;p&gt;FIGURE 4 AND ICA (R1). 
We agree that Fig. 4 should be improved to better appreciate p2 and p3 results. We will improve Fig. 4 with larger renderings and zoom-ins in the final version. We provided a supplementary video for 3D inspection. In Fig. 4, the ICA appears reduced in M_P because of a predefined ROI during CTA acquisition. The 3 clinical cases visualize only the LICA territory, since contrast was injected into the left ICA.&lt;/p&gt;

  &lt;p&gt;DIFFERENCE w.r.t. SOTA (R2): 
Our method differs from the population models (Cafaro, Zuo/Wu, Frisken) in that we use the patient’s own CTA as a prior, allowing us to reformulate biplane DSA reconstruction as a patient-specific 3D shape completion rather than inverse 3D reconstruction.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Invite for Rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes a CTA-guided Gaussian splatting framework for 3D cerebrovascular shape completion from biplane DSA.&lt;/p&gt;

      &lt;p&gt;Reviewers agree the problem could be clinically relevant and the approach is technically interesting, but raise concerns about limited and mostly qualitative clinical validation, reliance on simulated data, and unclear robustness to factors such as DSA synthesis and CTA-DSA registration. I suggest that the authors address these raised concerns in the rebuttal.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors provided a convincing rebuttal, and there is consensus among the reviewers that the paper presents a technically interesting and clinically motivated approach, which could be a valuable feasibility study. Therefore, I recommend acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-2&quot;&gt;Meta-review #2&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors have largely clarified the previously unclear points in their rebuttal, and the reviewers appear satisfied with the responses, with all agreeing on acceptance. The AC therefore recommends acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-3&quot;&gt;Meta-review #3&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;All reviewers agrees that the authors response were adequate.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Jehkul, Janik AND Frisken, Sarah AND Gopalakrishnan, Vivek AND Rueckert, Daniel AND Haouchine, Nazim</name></author><category term="Body -&gt; Vasculature" /><category term="Modalities -&gt; CT / X-ray" /><category term="Applications -&gt; Image-Guided Interventions" /><category term="Machine Learning -&gt; Deep Learning" /><category term="Jehkul, Janik" /><category term="Frisken, Sarah" /><category term="Gopalakrishnan, Vivek" /><category term="Rueckert, Daniel" /><category term="Haouchine, Nazim" /><summary type="html">Abstract We propose a novel approach to bridge the resolution and dimensionality gap between CTA, which provides 3D vascular geometry but often misses small vessels, and biplanar 2D DSA, which offers higher spatial resolution but lacks 3D structural information, by formulating the problem as a 3D shape completion task. Our method represents vasculature using a set of 3D Gaussians initialized from CTA-derived vessel geometry and augmented with additional spatial and opacity primitives seeded from two DSA projections. These Gaussians jointly encode geometry and attenuation and are optimized to fit the observed DSA images while remaining consistent with the original CTA anatomy. Experiments on synthetic and clinical cerebrovascular data demonstrate improved 3D reconstruction of small vessel branches at submillimetric resolution, validating the use of DSA to complement missing CTA vessel anatomy. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/2705_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: https://papers.miccai.org/miccai-2026/supp/2705_supp.zip Link to the Code Repository https://github.com/janik-j/cta-dsa-fusion Link to the Dataset(s) TopBrain CTA Dataset: https://topbrain2025.grand-challenge.org/data/ Clinical CTA/DSA dataset (Brigham and Women’s Hospital): Not publicly available due to PHI constraints BibTex @InProceedings{JehJan_3D_MICCAI2026,         author = { Jehkul, Janik AND Frisken, Sarah AND Gopalakrishnan, Vivek AND Rueckert, Daniel AND Haouchine, Nazim},         title = { { 3D Cerebrovascular Shape Completion from Biplane Angiography and CTA Prior } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16889},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper This work presents a Gaussian splatting-based method to reconstruct 3D mesh from bi-plane DSA, using the pre-operative CTA as the prior. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The proposed method seems to outperform the SOTA methods based on qualitative and quantitative assessments. 2.Loss functions are tailored to the task. 3.The ablation studies are comprehensive and clearly demonstrate the contribution of each component. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.Only a subset of TopBrain CTA data was used, but the authors didn’t provide the reason why they exclude the rest. 2.The DSA data were synthesised from CTA, which could be unreliable. In reality, CTA can visualise full cerebral vasculature whereas DSA only visualise part of it, because of the difference in the contrast injection methods (intravenous vs intra-arterial). In addition, there might be variation due to contrast agent dispersion. The authors should provide further information on how the DSA data synthesis was performed. 3.The method was tested on only three patients with acquired DSA and CTA. The projection of the reconstructed mesh seemingly agrees well with the original DSA in Patient 1.However, visual assessment based on the provided 3D meshes is very difficult for Patient 2 &amp;amp; 3 and unable to confirm successful reconstruction. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The authors claimed to release the source code and/or dataset upon acceptance of the submission. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html 1.The proposed method relies on CTA/DSA registration for real world application. How reliable is the registration used in this work and how significant would registration errors affect the final result? 2.The method employs multiple loss terms—could the authors elaborate on how their weights were balanced? Table 2 shows that combining all the designated losses does not achieve the best performance. Please provide further discussion. 3.(Figure 4) Why is the internal carotid artery missing in the incomplete mesh in every patient? As the main artery, it should be easily identified at initialisation. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? I am not convinced that the proposed method has achieved the claimed sucess. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The authors have addressed most of the previous concerns. Importantly, they have agreed to update Fig 4 to improve visualisation and to support the claims. While the small in vivo dataset of three patients makes the validation qualitative and speculative, it might be sufficient for proof of concept. Overall, the recommendation is acceptance. Review #2 Please describe the contribution of the paper The paper proposes a 3D cerebrovascular shape-completion method that combines an incomplete CTA-derived vessel prior with two biplane DSA projections. The core idea is to represent the vasculature using 3D Gaussians initialized from both the CTA mesh and DSA-derived residual context points, and then optimize them with a rendering loss, a false-negative vessel loss, and a prior-anchoring regularizer. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The paper addresses a clinically meaningful problem. CTA provides 3D structure but, as the authors stress, often misses small distal vessels, while biplane DSA has higher spatial resolution but only 2D information. Thus, framing this as shape completion from a patient-specific CTA prior plus biplane DSA is a strong and well-motivated formulation The method itself is technically interesting. The combination of dual-source Gaussian initialization, differentiable X-ray rasterization, a false-negative loss targeted at weakly supervised thin branches, and a prior-anchoring term is coherent and well aligned with the imaging physics and the stated goal of vessel completion rather than generic reconstruction. Overall, it feels like a quite interesting and novel approach Quantitative results are strong on the synthetic benchmark. The method clearly outperforms the reported baselines in both 2-view and 3-view settings The paper also includes meaningful ablation studies on losses, initialization strategy, and prior completeness The paper includes an initial clinical demonstration. Although a bit limited, the three-patient qualitative study supports the claim that the method can recover distal branches visible in DSA but absent from CTA Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. The main weakness is the limited real-world validation. Quantitative experiments are performed on only 10 synthetic CTA scans with simulated incomplete priors, while the clinical study covers only 3 patients and is qualitative only. There is no paired clinical ground truth or quantitative evaluation on real patient data, so it is still unclear how robust the method is under realistic acquisition artifacts, contrast inhomogeneity, and registration error A second weakness is that the novelty is somewhat incremental at the problem level, even if the specific combination is novel. Prior work has already studied two-view cerebral vascular reconstruction and learned priors for this setting (e.g., “Two Projections Suffice for Cerebral Vascular Reconstruction” by Frisken et al.). The novelty of the present paper lies more in the CTA-guided Gaussian splatting completion formulation than in introducing the two-view reconstruction problem itself A more minor weakness is that reproducibility is only partial in the current version. The paper gives the overall formulation and some hyperparameters, but several implementation details appear under-specified in the main text, such as some initialization and threshold choices for the auxiliary terms and prior neighborhood construction Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The authors claimed to release the source code and/or dataset upon acceptance of the submission. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html I think it’s an interesting paper overall and well-written. I liked the idea of using a CTA prior to somewhat turn biplane DSA reconstruction into a completion problem rather than a purely unconstrained inverse problem. The ablations are helpful and the improvements on the synthetic benchmark are convincing. The main issue for me is validation. The paper would be significantly stronger with either quantitative evaluation on clinical data or a more realistic simulation study that better captures registration error, residual subtraction artifacts, and contrast inhomogeneity. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? I am leaning weak accept because the paper presents a solid and clinically relevant methodological contribution: a well-motivated CTA-guided Gaussian-splatting formulation for completing cerebrovascular anatomy from only two DSA views. The quantitative gains over the reported baselines are substantial, and the ablations help support the design choices. At the same time, the evidence for real-world clinical utility is still limited. The quantitative study is small and synthetic, and the clinical evaluation is qualitative on only three patients. So to summarize, I find the idea strong and potentially impactful but I think the paper still sits near the threshold because the empirical validation does not yet fully match the ambition of the claims. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. I believe the authors have adequately addressed the concerns raised in my review. I am satisfied with the authors’ response. Review #3 Please describe the contribution of the paper Authors are proposing taking advantage of imaging methods with higher resolution to complement imaging methods with a lower resolution. In this case, completing the vasculature observed in CTA which in most hospital setups can go up to 0.5mm (or 0.83mm in the particular clinical evaluation of this paper), with the one observed in DSA acquisitions, the gold standard in many neurovascular diseases like stroke, which tend to go up to 0.2mm. Particularly helpful in treatment of small aneurysms, stroke and potentially calcification in distal vessels. Authors method to complete the vasculature uses dual initialization, one using a prior mesh from the CTA and DSA biplanar projections (improved with a third oblique view). Loss function includes a rendering component, a false negative vessel loss and a prior anchoring regularizer, to cover all the aspects of the method. Method is evaluated on both synthetic and clinical datasets (n=10 and n=3 respectively). Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. Complete reconstruction of distal vessels having just 3 or even two additional DSA planes and extend the baseline CTA is a great piece of work. Not only because of the extension of diagnostic tools, but also because particular neurovascular diseases whose gold standard diagnostic image is DSA will benefit from having complete distal information for treatment planning. The loss function covers important aspects of the proposed method: how good is the render, regularize on the priors and check false negative on distal branches. The term is well thought and parameters of the training objective are well defined (potentially tuned). False loss term makes sense since the original name and main method is the study on the contribution of sparse (biplanar) DSA, however is worth asking on the review side, is it the same parameter value when using the oblique view? Given the fact of how much it improves the results. Are the terms performing the same on the both biplanar and triplanar setups? Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. While the community broadly accepts demonstrations of clinical feasibility on small cohorts, the clinical evaluation (n=3) remains exclusively qualitative. Given that the paper explicitly claims clinical applicability and demonstrates recovery of 0.216mm vessels within 0.833mm CTA resolution, it would be interesting to disclose quantitative metrics for the clinical cases. Was the decision to keep this qualitative a consequence of annotation difficulty, or was there a deliberate methodological reason? Clarifying this would help readers assess the generalizability of the results. The method’s performance somehow relies on accurate 2D-3D registration between DSA planes and the CTA-derived mesh. However, no registration error analysis or sensitivity study is provided. Understanding there are space limitations and authors covered quite well the complete explanation, a small discussion on how robust the pipeline is to misalignment would be good. A small extension of the previous point is that you had space enough to cover a context point initialization ablation, however information of this table is limited. It is interesting why having less context points (only M1) perform better than uniform initialization. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The authors claimed to release the source code and/or dataset upon acceptance of the submission. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The paper presents a technically well explained method and a clinically motivated contribution at the using CTA and DSA imaging. The synthetic evaluation is well thought, the loss design is well motivated, and the gains from adding a third oblique plane are really notable. The contribution is novel and the clinical framing is, at least, relevant to diseases like stroke. The score is Weak Accept given the absence of quantitative metrics (CD, HD) in the clinical evaluation needs justification, if it was a deliberate design decision or an annotation constraint should be explicitly stated. Registration quality between DSA can be commented/expanded, as well as how brief was table 3 commented. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The rebuttal addresses my main concerns directly. The justification for the qualitative-only clinical evaluation, that quantitative ground truth would require CTA acquired at near-DSA resolution, which is not standard in clinical practice, is a methodologically honest answer, and the reframing of the 10 TopBrain cases as the quantitative cohort and the 3 clinical cases as feasibility-under-real-DSA conditions is reasonable, but still quite small, worth noting. The registration response (XVR with sub-mm accuracy, combined with Eq. 5’s KNN-anchored prior absorbing sub-voxel imprecision) explains the design feature that confers robustness, but not explain at all how good does registration need to be, a full sensitivity study would not hurt. The loss-weight clarification at biplane vs triplane answered my question about whether the loss terms perform equivalently across setups The contribution looks good and clinically motivated. Author Feedback We are glad reviewers found the problem “clinically meaningful” (R2); our method “novel”, “technically interesting” (R2), and “well explained” (R3) with a “strong and well-motivated” formulation (R2) and losses “tailored to the task” (R1). Reviews also found the results to “outperform[s] the SOTA methods” (R1) and the ablations to be “comprehensive” (R1) and that we overall presented a “great piece of work” (R3). We thank the reviewers for their comments and address the main ones below. CLINICAL VALIDATION (R1,R2,R3): We presented a quantitative evaluation on 10 TopBrain cases using real CTA and synthetic DSA, including comparisons with SOTA methods, an ablation study, and a sensitivity analysis, along with qualitative results on 3 clinical cases. These additional clinical cases are provided to demonstrate feasibility under real clinical DSA acquisition conditions, not to serve as a statistically powered evaluation cohort. Moreover, quantitative clinical evaluation would require CTA acquisition at a resolution comparable to DSA (to capture tiny vessels), which is not widely acquired in clinical practice. In addition, since our approach is a patient-specific optimization, all 10 cases contribute independently to the results in Tables 1-3, covering different anatomies and vascular variability. Altogether, while we agree that a larger clinical evaluation is beneficial, we believe that our evaluation supports the claims of the paper and is overall on par with related methodological papers in this area. SYNTHETIC DSA (R1): We agree CTA covers more vasculature than DSA. Our aim is not to match coverage but to use CTA as a prior to reach the sub-mm resolution that only DSA resolves, as stated in the introduction. We forward-project each CTA volume with TIGRE under a simulated cone-beam geometry and subtract mask from fill projections (Sec. 5.1). The clinical input I_v is a temporal MIP (Sec. 2.1), so both inputs reduce to a single static projection per view. R1 rightly notes that real DSA only depicts the injected territory due to IA vs IV contrast administration. The synthesis does not model this, but the clinical pipeline crops M_P to points projecting into DSA vessel masks across all views (Sec. 5.1.2), restricting the prior to the injected territory at input. We will clarify this in the final paper. TOPBRAIN SUBSET (R1, R2): We had to use a subset of TopBrain to fine-tune DenoiseNet, leaving 10 cases for evaluation. No method sees test data at training time. We will state this split explicitly in the paper. LOSS WEIGHTS (R1, R3): Each loss targets one error mode by design, so the full configuration is not best on every per-region metric. L_FN improves the missing region HD_ROI by 18%, L_PA the prior region CD_prior by 8%, combined best overall CD. λ_FN and λ_PA are identical at 2 and 3 views with effects more pronounced at 3 views (R3). λ_ssim differs (0 vs 0.07), as we empirically found that pure L1 works better at 2 views. REGISTRATION (R1, R3): CTA to DSA alignment uses XVR [7] (sub-mm accuracy). The CTA prior is not fixed: Eq. 5 anchors centers to the local mean of their KNN prior points, and the squared norm penalizes large displacements, so sub-voxel imprecision is absorbed without prior drift. FIGURE 4 AND ICA (R1). We agree that Fig. 4 should be improved to better appreciate p2 and p3 results. We will improve Fig. 4 with larger renderings and zoom-ins in the final version. We provided a supplementary video for 3D inspection. In Fig. 4, the ICA appears reduced in M_P because of a predefined ROI during CTA acquisition. The 3 clinical cases visualize only the LICA territory, since contrast was injected into the left ICA. DIFFERENCE w.r.t. SOTA (R2): Our method differs from the population models (Cafaro, Zuo/Wu, Frisken) in that we use the patient’s own CTA as a prior, allowing us to reformulate biplane DSA reconstruction as a patient-specific 3D shape completion rather than inverse 3D reconstruction. Meta-Review Meta-review #1 Your recommendation Invite for Rebuttal Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. This paper proposes a CTA-guided Gaussian splatting framework for 3D cerebrovascular shape completion from biplane DSA. Reviewers agree the problem could be clinically relevant and the approach is technically interesting, but raise concerns about limited and mostly qualitative clinical validation, reliance on simulated data, and unclear robustness to factors such as DSA synthesis and CTA-DSA registration. I suggest that the authors address these raised concerns in the rebuttal. After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. The authors provided a convincing rebuttal, and there is consensus among the reviewers that the paper presents a technically interesting and clinically motivated approach, which could be a valuable feasibility study. Therefore, I recommend acceptance. Meta-review #2 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. The authors have largely clarified the previously unclear points in their rebuttal, and the reviewers appear satisfied with the responses, with all agreeing on acceptance. The AC therefore recommends acceptance. Meta-review #3 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. All reviewers agrees that the authors response were adequate. back to top</summary></entry><entry><title type="html">3D Classification of Paramagnetic Rim Lesions in Multiple Sclerosis via Asymmetric QSM–FLAIR Modeling</title><link href="https://papers.miccai.org/miccai-2026/0002-Paper5533" rel="alternate" type="text/html" title="3D Classification of Paramagnetic Rim Lesions in Multiple Sclerosis via Asymmetric QSM–FLAIR Modeling" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0002-Paper5533</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0002-Paper5533">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;Paramagnetic rim lesions (Rim+) identified on susceptibility-sensitive MRI have recently emerged as a specific biomarker of chronic active inflammation in Multiple Sclerosis (MS) and are associated with long-term disability progression. However, susceptibility imaging and expert interpretation remain limited to specialized centers, visual assessment is time-consuming and variable, and the low prevalence of Rim+ lesions poses severe class imbalance challenges for automated analysis.
We propose FRODO, a 3D Fusion framework for Rim lesion classificatiOn using multimodal Deep-learning neurOimaging, designed for lesion-level Rim+/Rim- classification from Quantitative Susceptibility Mapping (QSM) and FLAIR MRI. 
The architecture explicitly models modality asymmetry by treating QSM as the primary susceptibility-driven signal and conditioning it with FLAIR-derived structural context. To improve robustness under limited data, we employ self-supervised multimodal pretraining followed by supervised fine-tuning with contrastive regularization.
The method was evaluated on a clinically acquired cohort of 88 people with MS with expert lesion annotations as reference standard. Results highlight improved performance compared to prior architectures, supporting the effectiveness of asymmetric multimodal modeling for automated chronic active lesion identification.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/5533_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/5533_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/veronicapignedoli/FRODO&quot;&gt;https://github.com/veronicapignedoli/FRODO&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{PigVer_3D_MICCAI2026,
        author = { Pignedoli, Veronica AND Boffa, Giacomo AND Noceti, Nicoletta AND Inglese, Matilde AND Odone, Francesca AND Moro, Matteo},
        title = { { 3D Classification of Paramagnetic Rim Lesions in Multiple Sclerosis via Asymmetric QSM–FLAIR Modeling } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16885},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes a 3D multimodal deep learning framework for lesion-level classification of paramagnetic rim lesions (Rim+) in multiple sclerosis using QSM and FLAIR MRI. The method introduces an asymmetric multimodal design that treats QSM as the primary susceptibility-driven signal and conditions it with FLAIR via spatial FiLM. To address limited data and strong class imbalance, the framework incorporates self-supervised cross-modal pretraining and supervised contrastive regularization.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ol&gt;
        &lt;li&gt;Well-motivated clinical problem. 
Automated identification of paramagnetic rim lesions is clinically important and challenging, with clear relevance for disease characterization and prognosis in multiple sclerosis.&lt;/li&gt;
        &lt;li&gt;Reasonable and task-aligned multimodal design. 
The asymmetric QSM–FLAIR formulation is well justified from a physiological perspective, and the use of FiLM-based conditioning provides an intuitive mechanism to incorporate structural context while preserving susceptibility-specific information.&lt;/li&gt;
        &lt;li&gt;Careful experimental setup. 
The use of a clinically acquired cohort, patient-level cross-validation, and appropriate evaluation metrics (particularly PR-AUC under strong class imbalance) strengthens the validity of the study.&lt;/li&gt;
        &lt;li&gt;Meaningful empirical improvements. 
The method shows clear gains over prior approaches, especially in PR-AUC and precision, which are important for rare lesion detection.&lt;/li&gt;
      &lt;/ol&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ol&gt;
        &lt;li&gt;Methodological novelty is moderate. 
The proposed framework mainly combines existing components, including 3D CNN backbones, FiLM conditioning, self-supervised pretraining, and contrastive learning. The contribution lies more in their integration for this specific task rather than in fundamentally new methodology.&lt;/li&gt;
        &lt;li&gt;Comparison to broader PRL literature is limited. 
The evaluation focuses primarily on QSMRim-Net and a ResNet baseline. Additional comparisons with other recent PRL detection/classification approaches (e. g. , studies such as PMC:12516162) would better contextualize the proposed method and strengthen claims of improved performance.&lt;/li&gt;
        &lt;li&gt;Limited generalization assessment. 
All experiments are conducted on a single-center cohort. While the evaluation protocol is well designed, external validation on independent datasets would be important to assess robustness across acquisition settings.&lt;/li&gt;
      &lt;/ol&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors claimed to release the source code and/or dataset upon acceptance of the submission.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(5) Accept — should be accepted, independent of rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Overall, this paper addresses a clinically relevant and technically challenging problem with a well-designed multimodal framework and solid experimental validation. The asymmetric modeling of QSM and FLAIR is intuitive and supported by meaningful improvements, particularly under severe class imbalance. While the methodological novelty is somewhat incremental and comparisons could be more comprehensive, the work provides practical value for automated PRL analysis. I would support acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents an end‑to‑end 3D multimodal deep learning framework for lesion‑level classification of paramagnetic rim (Rim+) versus non‑rim (Rim−) lesions in multiple sclerosis using QSM and FLAIR MRI. The method explicitly models modality asymmetry by leveraging QSM as the primary susceptibility‑driven signal and conditioning it with FLAIR‑based structural context, complemented by self‑supervised pretraining and contrastive regularization. The experimental results demonstrate improved Rim+ classification performance on a clinically representative dataset characterized by strong class imbalance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;The paper is clearly written, well structured, and easy to follow.&lt;/li&gt;
        &lt;li&gt;The proposed methodology is technically sound and well aligned with topics of interest to the MICCAI community.&lt;/li&gt;
        &lt;li&gt;The experimental evaluation suggests that the proposed approach can substantially improve Rim+ lesion classification performance compared to prior methods.&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;The cohort size used for training, validation, and testing is relatively small, which makes it difficult to assess the broader clinical impact and robustness of the proposed method.&lt;/li&gt;
        &lt;li&gt;The absence of a separately held, fully independent test set limits the assessment of generalizability across sites or imaging protocols.&lt;/li&gt;
        &lt;li&gt;In Table 1, the effect of the FLAIR modality is not clearly isolated. It would be helpful to clarify whether the same QSM/FLAIR input pairs were used consistently across QSMRim‑Net, ResNet, and the proposed method to ensure a fair comparison.&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;Given the limited amount of QSM data, generative modeling approaches for QSM‑based data augmentation could be discussed as a potential avenue to further mitigate data scarcity.&lt;/li&gt;
        &lt;li&gt;Section 2.1: The statement that “spatial conditioning enables localized contextual adaptation while preserving QSM as the dominant modality” would benefit from additional clarification or empirical justification.&lt;/li&gt;
        &lt;li&gt;Section 3.1: Please clarify whether both QSM and FLAIR images are consistently used at the lesion (connected‑component) level during classification.&lt;/li&gt;
        &lt;li&gt;For the self‑supervised pretraining stage, it would be useful to discuss whether larger unannotated datasets could have been leveraged to further address the limited dataset size, and if not, why this was not feasible.&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;My recommendation is based on the paper’s strong technical contribution and its successful application of multimodal deep learning to a clinically significant task. The authors present a well-motivated framework that effectively handles modality asymmetry and class imbalance in MS lesion classification. However, they reflect a need for greater clarity regarding experimental controls and concerns over the robustness of findings given the dataset constraints.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents an algorithm to perform binary classification of multiple sclerosis paramagnetic rim lesions (PRLs) on QSM and FLAIR MRI, with particular design considerations for multi-modal fusion and handling the challenges of a small data set and highly unbalanced classification task. The main benefit shown is increased positive predictive value (precision), while maintaining sensitivity, compared to the main comparator used in the study (QSMRim-Net).&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;PRLs are important to identify in multiple sclerosis, so the clinical need is real and topical.&lt;/li&gt;
        &lt;li&gt;The choice of QSM and FLAIR as the input sequences seems reasonable and practical for clinical use.&lt;/li&gt;
        &lt;li&gt;The included algorithmic components make logical sense. For example, the choice of FiLM seems to make sense for this application.&lt;/li&gt;
        &lt;li&gt;There is a good level of detail on the algorithm and training procedures.&lt;/li&gt;
        &lt;li&gt;The evaluation metrics are appropriate, particularly PR AUC.&lt;/li&gt;
        &lt;li&gt;An ablation study on the main components is included.&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;The data set is very small (88 people with MS), and due to heterogeneity of MS, is unlikely to be a “clinically representative” cohort as claimed. There is also no description of the clinical phenotypes (e.g., relapsing vs. progressive MS, etc.), which is important for assessing the suitability of the data.&lt;/li&gt;
        &lt;li&gt;The benefits of the algorithm shown in the paper may not hold for larger, real-world data sets.&lt;/li&gt;
        &lt;li&gt;Performance was only assessed at the lesion level, whereas patient-level classification is the most important task for predicting clinical progression.&lt;/li&gt;
        &lt;li&gt;While the proposed method is better than the main comparator, the sensitivity is still quite low (0.447), which is probably below clinical utility.&lt;/li&gt;
        &lt;li&gt;Statistical testing is only reported for ROC-AUC, which is not the most useful metric for highly unbalanced data.&lt;/li&gt;
        &lt;li&gt;One key reference (which also uses QSM and FLAIR), seems to be missing: https://arxiv.org/pdf/2303.08434&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors claimed to release the source code and/or dataset upon acceptance of the submission.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The algorithm seems well thought-out, but the small data set, selective statistical testing, and missing reference make the overall validation fairly light.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;We thank the Meta-Reviewer and the Reviewers for their constructive comments, which helped us clarify and further develop key aspects of the manuscript. We have carefully addressed the minor comments, including improving explanations where requested and adding missing literature citations. Specifically, we have expanded the description of the clinical phenotypes and revised the Introduction to incorporate the additional references highlighted by the Reviewers. This includes adding the PRL literature mentioned by Reviewer #1 (https://pmc.ncbi.nlm.nih.gov/articles/PMC12516162/) and replacing the arXiv reference noted by Reviewer #3 with the official MICCAI 2023 publication (https://link.springer.com/chapter/10.1007/978-3-031-43895-0_72), as also suggested by the Meta-Reviewer.
Furthermore, in the camera-ready manuscript, we addressed the following reviewers’ requests for clarification.
-) Methodological Clarifications. As requested by Reviewer #2, we clarified that the same QSM/FLAIR input pairs were consistently used across QSMRim-Net, ResNet, and our proposed method to ensure a fair comparison in Table 1.We also explicitly state that both modalities are jointly used at the lesion level during classification.
-) Spatial Conditioning. We expanded the explanation in Section 2.1 to better justify how spatial conditioning enables localized contextual adaptation while preserving QSM as the dominant modality.
-) Extended Discussions. We added a discussion on potential future directions to mitigate data scarcity, including the use of generative modeling for data augmentation and the feasibility of leveraging larger unannotated datasets for self-supervised pretraining.
-) Study Limitations. As pointed out by all Reviewers and the Meta-Reviewer, we acknowledged the limited size of our single-center cohort (88 people) that may limit the broader generalization of our findings. Indeed, as stated in the Conclusion Section of our camera-ready manuscript, our primary future direction is to validate the proposed framework across larger, multi-center datasets to ensure robustness to acquisition variability. 
-) Statistical and Patient-Level Analysis. We incorporated a more comprehensive statistical analysis of the PR-AUC metric to better reflect performance on highly imbalanced data, and we added a discussion of patient-level results, as suggested by Reviewer #3.In particular, building on the considerations reported in [https://pubmed.ncbi.nlm.nih.gov/35247730/] and adapting them to our dataset, we classify a participant as Rim-positive if the model predicts at least one Rim+ lesion. In particular, we classify a participant as Rim-positive if the model predicts at least one Rim+ lesion. Under this definition, our model yields a more balanced patient-level diagnostic profile and substantially outperforms the baseline in estimating the total chronic active lesion burden, achieving a higher Pearson correlation between predicted and ground-truth lesion counts (r = 0.755 vs. r = 0.439). To assess statistical significance for PR-AUC, we employed a 1000-replicate bootstrapping procedure on out-of-fold predictions. The resulting 95% confidence intervals confirm that the improvement of our model is statistically significant, as its interval strictly exceeds and does not overlap with that of QSMRim-Net.
As suggested by the Meta-Reviewer, we did not conduct additional experiments; however, we believe these revisions significantly strengthen the context, transparency, and clarity of our work.&lt;/p&gt;

&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Provisional Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This work presents a solution to an important clinical problem, with a clinically sound methodology and a well-written manuscript. These strengths are highlighted by all reviewers. I would also like to highlight the overall high scores. That is the basis of my recommendation.&lt;/p&gt;

      &lt;p&gt;Having said that, there are concerns (some of them shared by all reviewers) that I think should be addressed on a final manuscript. In the following sentences I highlight what I consider important issues to address, even though I would also recommend to check all the feedback from the reviewers.&lt;/p&gt;

      &lt;p&gt;While I want to emphasise that no new experiments should be performed, I believe that the current results should be extended following reviewer suggestions. Important things to include would be an explanation of why some studies were missed and their inclusion in the introduction [R#1 and #3], acknowledgment of the limited impact of the results and generalisation due to a single cohort and low sensitivity [R#1, 2 and 3], a more thorough statistical analysis for the non-AUC metrics [R#3] and results at the patient level [R#3]. Regarding the publication mentioned by reviewer #3, I would also advise the authors to use the official publication (MICCAI 2023 paper: https://link.springer.com/chapter/10.1007/978-3-031-43895-0_72) and not the pre-print (arxiv link).&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Pignedoli, Veronica AND Boffa, Giacomo AND Noceti, Nicoletta AND Inglese, Matilde AND Odone, Francesca AND Moro, Matteo</name></author><category term="Body -&gt; Brain" /><category term="Modalities -&gt; MRI" /><category term="Applications -&gt; Anomaly / Lesion Detection" /><category term="Applications -&gt; Computer-Aided Diagnosis" /><category term="Machine Learning -&gt; Deep Learning" /><category term="Machine Learning -&gt; Multimodal Models / LLMs / VLMs" /><category term="Pignedoli, Veronica" /><category term="Boffa, Giacomo" /><category term="Noceti, Nicoletta" /><category term="Inglese, Matilde" /><category term="Odone, Francesca" /><category term="Moro, Matteo" /><summary type="html">Abstract Paramagnetic rim lesions (Rim+) identified on susceptibility-sensitive MRI have recently emerged as a specific biomarker of chronic active inflammation in Multiple Sclerosis (MS) and are associated with long-term disability progression. However, susceptibility imaging and expert interpretation remain limited to specialized centers, visual assessment is time-consuming and variable, and the low prevalence of Rim+ lesions poses severe class imbalance challenges for automated analysis. We propose FRODO, a 3D Fusion framework for Rim lesion classificatiOn using multimodal Deep-learning neurOimaging, designed for lesion-level Rim+/Rim- classification from Quantitative Susceptibility Mapping (QSM) and FLAIR MRI. The architecture explicitly models modality asymmetry by treating QSM as the primary susceptibility-driven signal and conditioning it with FLAIR-derived structural context. To improve robustness under limited data, we employ self-supervised multimodal pretraining followed by supervised fine-tuning with contrastive regularization. The method was evaluated on a clinically acquired cohort of 88 people with MS with expert lesion annotations as reference standard. Results highlight improved performance compared to prior architectures, supporting the effectiveness of asymmetric multimodal modeling for automated chronic active lesion identification. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/5533_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository https://github.com/veronicapignedoli/FRODO Link to the Dataset(s) N/A BibTex @InProceedings{PigVer_3D_MICCAI2026,         author = { Pignedoli, Veronica AND Boffa, Giacomo AND Noceti, Nicoletta AND Inglese, Matilde AND Odone, Francesca AND Moro, Matteo},         title = { { 3D Classification of Paramagnetic Rim Lesions in Multiple Sclerosis via Asymmetric QSM–FLAIR Modeling } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16885},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper This paper proposes a 3D multimodal deep learning framework for lesion-level classification of paramagnetic rim lesions (Rim+) in multiple sclerosis using QSM and FLAIR MRI. The method introduces an asymmetric multimodal design that treats QSM as the primary susceptibility-driven signal and conditions it with FLAIR via spatial FiLM. To address limited data and strong class imbalance, the framework incorporates self-supervised cross-modal pretraining and supervised contrastive regularization. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. Well-motivated clinical problem. Automated identification of paramagnetic rim lesions is clinically important and challenging, with clear relevance for disease characterization and prognosis in multiple sclerosis. Reasonable and task-aligned multimodal design. The asymmetric QSM–FLAIR formulation is well justified from a physiological perspective, and the use of FiLM-based conditioning provides an intuitive mechanism to incorporate structural context while preserving susceptibility-specific information. Careful experimental setup. The use of a clinically acquired cohort, patient-level cross-validation, and appropriate evaluation metrics (particularly PR-AUC under strong class imbalance) strengthens the validity of the study. Meaningful empirical improvements. The method shows clear gains over prior approaches, especially in PR-AUC and precision, which are important for rare lesion detection. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. Methodological novelty is moderate. The proposed framework mainly combines existing components, including 3D CNN backbones, FiLM conditioning, self-supervised pretraining, and contrastive learning. The contribution lies more in their integration for this specific task rather than in fundamentally new methodology. Comparison to broader PRL literature is limited. The evaluation focuses primarily on QSMRim-Net and a ResNet baseline. Additional comparisons with other recent PRL detection/classification approaches (e. g. , studies such as PMC:12516162) would better contextualize the proposed method and strengthen claims of improved performance. Limited generalization assessment. All experiments are conducted on a single-center cohort. While the evaluation protocol is well designed, external validation on independent datasets would be important to assess robustness across acquisition settings. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The authors claimed to release the source code and/or dataset upon acceptance of the submission. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (5) Accept — should be accepted, independent of rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? Overall, this paper addresses a clinically relevant and technically challenging problem with a well-designed multimodal framework and solid experimental validation. The asymmetric modeling of QSM and FLAIR is intuitive and supported by meaningful improvements, particularly under severe class imbalance. While the methodological novelty is somewhat incremental and comparisons could be more comprehensive, the work provides practical value for automated PRL analysis. I would support acceptance. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #2 Please describe the contribution of the paper This paper presents an end‑to‑end 3D multimodal deep learning framework for lesion‑level classification of paramagnetic rim (Rim+) versus non‑rim (Rim−) lesions in multiple sclerosis using QSM and FLAIR MRI. The method explicitly models modality asymmetry by leveraging QSM as the primary susceptibility‑driven signal and conditioning it with FLAIR‑based structural context, complemented by self‑supervised pretraining and contrastive regularization. The experimental results demonstrate improved Rim+ classification performance on a clinically representative dataset characterized by strong class imbalance. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The paper is clearly written, well structured, and easy to follow. The proposed methodology is technically sound and well aligned with topics of interest to the MICCAI community. The experimental evaluation suggests that the proposed approach can substantially improve Rim+ lesion classification performance compared to prior methods. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. The cohort size used for training, validation, and testing is relatively small, which makes it difficult to assess the broader clinical impact and robustness of the proposed method. The absence of a separately held, fully independent test set limits the assessment of generalizability across sites or imaging protocols. In Table 1, the effect of the FLAIR modality is not clearly isolated. It would be helpful to clarify whether the same QSM/FLAIR input pairs were used consistently across QSMRim‑Net, ResNet, and the proposed method to ensure a fair comparison. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html Given the limited amount of QSM data, generative modeling approaches for QSM‑based data augmentation could be discussed as a potential avenue to further mitigate data scarcity. Section 2.1: The statement that “spatial conditioning enables localized contextual adaptation while preserving QSM as the dominant modality” would benefit from additional clarification or empirical justification. Section 3.1: Please clarify whether both QSM and FLAIR images are consistently used at the lesion (connected‑component) level during classification. For the self‑supervised pretraining stage, it would be useful to discuss whether larger unannotated datasets could have been leveraged to further address the limited dataset size, and if not, why this was not feasible. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? My recommendation is based on the paper’s strong technical contribution and its successful application of multimodal deep learning to a clinically significant task. The authors present a well-motivated framework that effectively handles modality asymmetry and class imbalance in MS lesion classification. However, they reflect a need for greater clarity regarding experimental controls and concerns over the robustness of findings given the dataset constraints. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #3 Please describe the contribution of the paper This paper presents an algorithm to perform binary classification of multiple sclerosis paramagnetic rim lesions (PRLs) on QSM and FLAIR MRI, with particular design considerations for multi-modal fusion and handling the challenges of a small data set and highly unbalanced classification task. The main benefit shown is increased positive predictive value (precision), while maintaining sensitivity, compared to the main comparator used in the study (QSMRim-Net). Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. PRLs are important to identify in multiple sclerosis, so the clinical need is real and topical. The choice of QSM and FLAIR as the input sequences seems reasonable and practical for clinical use. The included algorithmic components make logical sense. For example, the choice of FiLM seems to make sense for this application. There is a good level of detail on the algorithm and training procedures. The evaluation metrics are appropriate, particularly PR AUC. An ablation study on the main components is included. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. The data set is very small (88 people with MS), and due to heterogeneity of MS, is unlikely to be a “clinically representative” cohort as claimed. There is also no description of the clinical phenotypes (e.g., relapsing vs. progressive MS, etc.), which is important for assessing the suitability of the data. The benefits of the algorithm shown in the paper may not hold for larger, real-world data sets. Performance was only assessed at the lesion level, whereas patient-level classification is the most important task for predicting clinical progression. While the proposed method is better than the main comparator, the sensitivity is still quite low (0.447), which is probably below clinical utility. Statistical testing is only reported for ROC-AUC, which is not the most useful metric for highly unbalanced data. One key reference (which also uses QSM and FLAIR), seems to be missing: https://arxiv.org/pdf/2303.08434 Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The authors claimed to release the source code and/or dataset upon acceptance of the submission. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The algorithm seems well thought-out, but the small data set, selective statistical testing, and missing reference make the overall validation fairly light. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Author Feedback We thank the Meta-Reviewer and the Reviewers for their constructive comments, which helped us clarify and further develop key aspects of the manuscript. We have carefully addressed the minor comments, including improving explanations where requested and adding missing literature citations. Specifically, we have expanded the description of the clinical phenotypes and revised the Introduction to incorporate the additional references highlighted by the Reviewers. This includes adding the PRL literature mentioned by Reviewer #1 (https://pmc.ncbi.nlm.nih.gov/articles/PMC12516162/) and replacing the arXiv reference noted by Reviewer #3 with the official MICCAI 2023 publication (https://link.springer.com/chapter/10.1007/978-3-031-43895-0_72), as also suggested by the Meta-Reviewer. Furthermore, in the camera-ready manuscript, we addressed the following reviewers’ requests for clarification. -) Methodological Clarifications. As requested by Reviewer #2, we clarified that the same QSM/FLAIR input pairs were consistently used across QSMRim-Net, ResNet, and our proposed method to ensure a fair comparison in Table 1.We also explicitly state that both modalities are jointly used at the lesion level during classification. -) Spatial Conditioning. We expanded the explanation in Section 2.1 to better justify how spatial conditioning enables localized contextual adaptation while preserving QSM as the dominant modality. -) Extended Discussions. We added a discussion on potential future directions to mitigate data scarcity, including the use of generative modeling for data augmentation and the feasibility of leveraging larger unannotated datasets for self-supervised pretraining. -) Study Limitations. As pointed out by all Reviewers and the Meta-Reviewer, we acknowledged the limited size of our single-center cohort (88 people) that may limit the broader generalization of our findings. Indeed, as stated in the Conclusion Section of our camera-ready manuscript, our primary future direction is to validate the proposed framework across larger, multi-center datasets to ensure robustness to acquisition variability. -) Statistical and Patient-Level Analysis. We incorporated a more comprehensive statistical analysis of the PR-AUC metric to better reflect performance on highly imbalanced data, and we added a discussion of patient-level results, as suggested by Reviewer #3.In particular, building on the considerations reported in [https://pubmed.ncbi.nlm.nih.gov/35247730/] and adapting them to our dataset, we classify a participant as Rim-positive if the model predicts at least one Rim+ lesion. In particular, we classify a participant as Rim-positive if the model predicts at least one Rim+ lesion. Under this definition, our model yields a more balanced patient-level diagnostic profile and substantially outperforms the baseline in estimating the total chronic active lesion burden, achieving a higher Pearson correlation between predicted and ground-truth lesion counts (r = 0.755 vs. r = 0.439). To assess statistical significance for PR-AUC, we employed a 1000-replicate bootstrapping procedure on out-of-fold predictions. The resulting 95% confidence intervals confirm that the improvement of our model is statistically significant, as its interval strictly exceeds and does not overlap with that of QSMRim-Net. As suggested by the Meta-Reviewer, we did not conduct additional experiments; however, we believe these revisions significantly strengthen the context, transparency, and clarity of our work. Meta-Review Meta-review #1 Your recommendation Provisional Accept Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. This work presents a solution to an important clinical problem, with a clinically sound methodology and a well-written manuscript. These strengths are highlighted by all reviewers. I would also like to highlight the overall high scores. That is the basis of my recommendation. Having said that, there are concerns (some of them shared by all reviewers) that I think should be addressed on a final manuscript. In the following sentences I highlight what I consider important issues to address, even though I would also recommend to check all the feedback from the reviewers. While I want to emphasise that no new experiments should be performed, I believe that the current results should be extended following reviewer suggestions. Important things to include would be an explanation of why some studies were missed and their inclusion in the introduction [R#1 and #3], acknowledgment of the limited impact of the results and generalisation due to a single cohort and low sensitivity [R#1, 2 and 3], a more thorough statistical analysis for the non-AUC metrics [R#3] and results at the patient level [R#3]. Regarding the publication mentioned by reviewer #3, I would also advise the authors to use the official publication (MICCAI 2023 paper: https://link.springer.com/chapter/10.1007/978-3-031-43895-0_72) and not the pre-print (arxiv link). back to top</summary></entry><entry><title type="html">A 3D Unrolling Framework for Joint Sodium MRI Reconstruction and Concentration Quantification</title><link href="https://papers.miccai.org/miccai-2026/0003-Paper4553" rel="alternate" type="text/html" title="A 3D Unrolling Framework for Joint Sodium MRI Reconstruction and Concentration Quantification" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0003-Paper4553</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0003-Paper4553">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;Sodium MRI can non-invasively measure tissue sodium concentration (TSC), a key biomarker for stroke, tumor, cartilage degeneration, and other diseases, but severely suffers from intrinsically low signal-to-noise ratio and long scan times. Existing methods typically perform denoising on conventionally reconstructed images and then conduct TSC quantification separately, leading to oversmooth reconstruction for highly accelerated acquisition and prohibiting end-to-end concentration quantification. To fill this gap, we propose the first deep unrolling framework in sodium MRI, termed SoReCon, that takes gridded radial k-space as input and simultaneously performs sodium MRI reconstruction and TSC quantification. To this end, we for the first time craft a sharpness-enhanced training data tailored to obtaining vendor-style sodium MRI reconstruction, such that the reconstructed image achieves an appearance consistent with the vendor-reconstructed image without requiring any post-processing. Besides, we devise a differentiable concentration mapping module containing B1 inhomogeneity correction, automatic phantom detection, and linear transformation to convert the reconstructed image into a TSC map. Experiments on 107 subjects show that our method outperforms the existing representative methods by 7.19 dB in PSNR for image reconstruction and 2.54 mM in MAE for TSC quantification. Prospective evaluation on three additional subjects demonstrates the feasibility of SoReCon under real-world four-fold acceleration, reducing the primary sodium acquisition time from 15 min to 3.75 min and highlighting its potential for clinical translation.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/4553_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/4553_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{CaoYil_A3D_MICCAI2026,
        author = { Cao, Yilin AND Duan, Caohui AND Shen, Dinggang AND Lou, Xin AND Sun, Kaicong},
        title = { { A 3D Unrolling Framework for Joint Sodium MRI Reconstruction and Concentration Quantification } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16888},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors present a deep-learning based reconstruction method for accelerated sodium MRI. The networks enable end-to-end total sodium concentration (TSC) mapping from raw k-space data. In-vivo studies show that the proposed method preserves both reconstruction fidelity and quantitative accuracy at a fourfold acceleration factor.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The proposed method directly produces TSC images, which is of most interest in sodium MRI. In the current practice, deriving TSC from sodium MRI relies on simple (linear model) but tedious post-processing, often including manual masking for calibration phantoms. This work achieve automation for the post-processing step. 
2.This work also implemented bilateral filter-based reconstruction for use as the reference. The reference seemingly suppresses noise while preserving edges. 
3.Prospective study shows that the proposed method outperformed scanner’s reconstruction pipeline.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.p4: What is “clinical DICOM reconstruction”? Is it named after the image format? This is a very confusing nomenclature. 
2.The results are unconvincing and the mechanisms of the proposed method is not fully explored. Please refer to the comments below.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The authors claim that non-uniform FFT (NUFFT) reconstruction produced blur images. However, NUFFT as a Fourier transform variant does not inherently cause blurring (c. f. Behl et. al. , 2016, Magnetic Resonance in Medicine for clear 23Na images by NUFFT). Instead, the blurring might be rooted from density compensation. Authors should describe how the density compensation weighting was performed during NUFFT. Additionally, the acronym should be defined at its first occurrence in the manuscript. 
2.The authors should explain how the retrospective data set was divided into train/test sets. 
3.The effect of extra B1 correction was not fully discussed. (1) Please introduce NAA/NAV with a bit more details. (2) How much time does the extra scan cost? 
4.The CMM is not justified based on the presented evidence. The ablation study shows that the proposed CMM leads to lower PSNR and SSIM. Also, the authors should provide the performance in the case of 2x acceleration. 
5.Fig 5: It would be interesting to compare the proposed method compared with the bilateral filtered reconstruction on the undersampled k-space data.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;While this work presents an interesting method, the experiment design is probably flawed and there lacks critical evidence supporting authors’ claim.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Reject&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors have provided responses to questions regarding the soundness of results. However, I am afraid extra work is required to understand and verify the proposed design, particularly:
(1) the improvment seen after adding CMM and the effect of incorporating data from prolonged auxilliary scans;
(2) seemingly incoherent ablation study results;
(3) unexplained lower performance on prospective data when compared to the test set.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The main contributions include: 1) The first end-to-end framework from undersampled k-space to quantitative sodium maps by jointly performing sodium MRI reconstruction and sodium concentration quantification; 2) Construction of a k-space training set for sodium MRI reconstruction with enhanced sharpness; 3) Integrating a concentration mapping module and concentration loss into network training to enable concentration error backpropagation; 4) Evaluated on a relatively large in-house data containing 110 patients including a prospective study, reducing scan time from 15 minutes to 3.75 minutes (4×) while preserving reconstruction quality.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents several notable strengths. First, it addresses a clinically important problem in sodium MRI by tackling both low SNR and long acquisition time through a unified framework. Second, the proposed end-to-end approach that jointly performs image reconstruction and concentration quantification from undersampled k-space is well-motivated and represents a meaningful departure from conventional two-stage pipelines. Third, the integration of a differentiable concentration mapping module incorporating B1 correction and phantom calibration enhances the physiological relevance of the method. In addition, the study is supported by a relatively large in-house dataset and includes prospective validation, which strengthens its practical significance. Finally, the demonstrated acceleration with preserved reconstruction fidelity highlights the potential for real clinical translation.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The limited prospective validation restricts evidence of generalizability. Additionally, the ablation study does not fully justify the contribution of individual components.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(5) Accept — should be accepted, independent of rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The first end-to-end framework from undersampled k-space to quantitative sodium maps; Evaluated on a relatively large in-house data; A prospective study, demonstrating the reduction of scan time from 15 minutes to 3.75 minutes.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The first end-to-end framework from undersampled k-space to quantitative sodium maps; Evaluated on a relatively large in-house data; A prospective study, demonstrating the reduction of scan time from 15 minutes to 3.75 minutes.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper proposes a new learning technique for 23Na reconstruction.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;*The approach is new
*Evaluations are performed with real prospective data&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;*The paper does not compare against reasonable baseline techniques.&lt;/p&gt;

      &lt;p&gt;*The paper focuses on PSNR, SSIM, and MAE without identifying limitations. This approach has been recently criticized. (DOI:10.1002/mrm.70377)&lt;/p&gt;

      &lt;p&gt;*The literature review has a major omission with respect to sodium reconstruction methods that use 1H reference images for denoising and edge preservation:&lt;/p&gt;

      &lt;p&gt;Atkinson IC, Thulborn KR, Lu A, Haldar J, Zhou XJ, Claiborne T, Liang ZP. Quantitative 23-sodium and 17-oxygen MR imaging in human brain at 9.4 Tesla enhanced by constrained k-space reconstruction. In Proceedings of the 16th Annual Meeting of ISMRM, Toronto 2008 (p. 335).&lt;/p&gt;

      &lt;p&gt;Gnahm, C., Bock, M., Bachert, P., Semmler, W., Behl, N.G.R. and Nagel, A.M. (2014), Iterative 3D projection reconstruction of 23Na data with an 1H MRI constraint. Magn. Reson. Med., 71: 1720-1732.
Gnahm C, Nagel AM. Anatomically weighted second-order total variation reconstruction of 23Na MRI using prior information from 1H MRI. Neuroimage. 2015 Jan 15;105:452-61.
Lachner S, Zaric O, Utzschneider M, Minarikova L, Zbýň Š, Hensel B, Trattnig S, Uder M, Nagel AM. Compressed sensing reconstruction of 7 Tesla 23Na multi-channel breast data using 1H MRI constraint. Magnetic resonance imaging. 2019 Jul 1;60:145-56.
Zhao Y, Guo R, Li Y, Thulborn KR, Liang ZP. High-resolution sodium imaging using anatomical and sparsity constraints for denoising and recovery of novel features. Magnetic Resonance in Medicine. 2021 Aug;86(2):625-36.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The novelty is borderline, the results are hard to interpret, and the paper has cited almost none of the relevant 23Na reconstruction methods.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors responded well to comments&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;We thank the meta-reviewer and reviewers for your constructive feedback and positive recognition of our work on end-to-end sodium MRI reconstruction and quantitative concentration mapping. Below we clarify the major concerns.&lt;/p&gt;

  &lt;p&gt;(1) Novelty, related work, and baseline positioning (R1, R3, MR).
We thank R3 for highlighting prior 1H-constrained methods (e.g., Gnahm et al.), which use external anatomical priors. Since our method relies on multi-coil k-space data for reconstruction rather than on T1w priors for postprocessing (denoising), our protocol did not acquire 1H data and we cannot use these 1H-constrained methods for comparison. Instead, we selected three baselines: NUFFT, compressed sensing, and 3D U-Net, covering both traditional and learning-based approaches and allowing fair comparison. We will add the 1H-constrained methods in the related work. We thank R1 for suggesting the bilateral-filtered reconstruction and will add it into comparison in our journal version.&lt;/p&gt;

  &lt;p&gt;(2) Clarification of NUFFT/regridding and terminology (R1, MR).
We apologize for the previous unclarity. For NUFFT, we used Pipe-Menon algorithm for density compensation and chose the best parameters (e.g., iteration number) for the highest PSNR/SSIM, yielding relatively blurred reconstruction. For training data preparation, we used regridding on raw data without density compensation and took it as input for our sharpening procedure, which balances noise suppression and structural visibility. We will release the code upon acceptance. We agree that “clinical DICOM reconstruction” is confusing and will replace it with “scanner-provided reconstruction”. We will clarify the above issues in revised version.&lt;/p&gt;

  &lt;p&gt;(3) Justification of CMM and evaluation metrics (R1, R2, R3, MR).
Our end-to-end framework aims to obtain both accurate sodium quantification and reconstruction fidelity. At 4× acceleration, CMM substantially improves TSC MAE (9.17→6.16 mM) with slight PSNR/SSIM changes, indicating 1) TSC is easier to recover than image reconstruction for high acceleration, and the model balances both; 2) reconstruction improvement does not always align with TSC [Haldar et al.,2026]. Both show the importance of end-to-end learning. At 2× acceleration, CMM improves both MAE (7.79→6.20 mM) and PSNR (39.55→40.16 dB). We did not show the results due to page limits. 
Besides PSNR/SSIM, we included TSC error, prospective validation, and visual comparison. Prospective images were reviewed by experienced radiologists and rated as excellent compared to scanner-provided images, showing its clinical potential. We will discuss metric limitations and refer to [Haldar et al.,2026] in the revised manuscript.&lt;/p&gt;

  &lt;p&gt;(4) Experimental protocol and B1 correction (R1, MR).
The retrospective dataset contains 110 cases. 3 were excluded due to quality control. The remaining 107 cases were split at subject level into 77/10/20 for training/validation/test.
For B1 calibration, due to B1+ transmit field inhomogeneity, image intensity may not accurately reflect the underlying sodium concentration. Therefore, we acquired high-SNR NAA and homogeneous NAV acquisitions and the voxel-wise NAA/NAV ratio estimates the relative transmit sensitivity map for quantitative bias correction. Our protocol includes two NAA scans for improved SNR (1 min each) and one NAV scan (6 min). These scans are part of our standard workflow. We will add these details to the revised version.&lt;/p&gt;

  &lt;p&gt;(5) Prospective validation and practical feasibility (R2, MR).
We acknowledge that the prospective data is limited due to high cost and long acquisition time. Nevertheless, we performed quantitative and visual evaluation on these prospective acquisitions and will collect more data in future to strengthen the clinical utility. From our empirical results, we found end-to-end learning brings the largest gain and CMM the second. Due to page limits, this was omitted but will be added in the revised version.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Invite for Rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;There seems to be concerns around interpretability of the results, innovation, and literature review, resulting in mixed reviews.&lt;/p&gt;

      &lt;p&gt;Strengths: The work addresses an important sodium MRI problem by proposing an end-to-end 3D unrolled framework that jointly reconstructs undersampled k-space and produces total sodium concentration maps, incorporating B1 correction, phantom calibration, a relatively large in-house dataset, and prospective validation. Weaknesses: Reviewers found the experimental evidence uneven, with limited prospective validation, incomplete ablation of components, unclear or possibly flawed experimental design, missing comparisons to strong or relevant sodium reconstruction baselines, and an insufficient literature review that makes the novelty hard to assess. The authors therefore be invited for the rebuttal.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors are advised to address the reviewers’ remaining concerns and do what they promise to do in their rebuttal.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-2&quot;&gt;Meta-review #2&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes a 3D unrolling framework for joint sodium MRI reconstruction and concentration quantification, aiming to reconstruct accelerated sodium MRI from undersampled k-space while directly producing quantitative total sodium concentration maps.&lt;/p&gt;

      &lt;p&gt;The reviewers recognized the clinical importance of accelerated sodium MRI, the value of an end-to-end framework from k-space to quantitative sodium maps, and the inclusion of B1 correction, phantom calibration, a relatively large in-house dataset, and prospective validation. They also appreciated the potential reduction of scan time from 15 minutes to 3.75 minutes.&lt;/p&gt;

      &lt;p&gt;The reviewers raised concerns about unclear experimental design, limited prospective validation, incomplete ablation, missing sodium MRI baseline comparisons, and insufficient discussion of related 1H-constrained sodium reconstruction methods. In the rebuttal, the authors clarified the NUFFT/regridding setup, dataset split, B1 calibration protocol, the role of the concentration mapping module, and the rationale for selected baselines. They also committed to improving the related-work discussion and terminology.&lt;/p&gt;

      &lt;p&gt;Although some limitations remain, especially the limited prospective data and the need for clearer ablation presentation, the rebuttal sufficiently addressed the main concerns, and most reviewers supported acceptance after rebuttal. Therefore, I recommend acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-3&quot;&gt;Meta-review #3&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Post-rebuttal consensus is 2 accept and 1 reject, with R3 moving from weak reject to weak accept. The paper is not a clear accept because R1’s technical objections remain meaningful, but the rebuttal improved the balance. I would recommend accept, emphasizing that the final version must include the promised clarifications, related work, metric limitations, and prospective validation caveats.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-4&quot;&gt;Meta-review #4&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper rebuilds undersampled sodium MRI and outputs tissue sodium concentration maps in one trained pass. The main worry at review was that the ablation made no sense, since the concentration module lowers PSNR and SSIM. The rebuttal clears this up.
The panel’s reconstruction expert read the same point and moved from reject to accept. The rest holds up from the submitted paper. There are three real baselines, and the extra methods one reviewer asked for need 1H scans this study never collects. PSNR and SSIM are weak metrics, but they are not the headline. The real number is concentration error in mm, which is the right thing to measure and it improves.&lt;/p&gt;

      &lt;p&gt;The work also has real clinical weight, with a large in-house dataset and a prospective study rated by radiologists. What is left is fixable at camera-ready: the 3.75 minute figure leaves out about 8 minutes of calibration scans and should be given as total time, the prospective set is weaker than the test set and needs a line of explanation, and the “first” claim is only true because it is limited to sodium, since the joint reconstruct-and-quantify idea itself comes from proton qMRI.&lt;/p&gt;

      &lt;p&gt;It is a sound, clinically useful paper at borderline accept bar.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Cao, Yilin AND Duan, Caohui AND Shen, Dinggang AND Lou, Xin AND Sun, Kaicong</name></author><category term="Body -&gt; Brain" /><category term="Modalities -&gt; MRI" /><category term="Applications -&gt; Image Reconstruction" /><category term="Machine Learning -&gt; Deep Learning" /><category term="Cao, Yilin" /><category term="Duan, Caohui" /><category term="Shen, Dinggang" /><category term="Lou, Xin" /><category term="Sun, Kaicong" /><summary type="html">Abstract Sodium MRI can non-invasively measure tissue sodium concentration (TSC), a key biomarker for stroke, tumor, cartilage degeneration, and other diseases, but severely suffers from intrinsically low signal-to-noise ratio and long scan times. Existing methods typically perform denoising on conventionally reconstructed images and then conduct TSC quantification separately, leading to oversmooth reconstruction for highly accelerated acquisition and prohibiting end-to-end concentration quantification. To fill this gap, we propose the first deep unrolling framework in sodium MRI, termed SoReCon, that takes gridded radial k-space as input and simultaneously performs sodium MRI reconstruction and TSC quantification. To this end, we for the first time craft a sharpness-enhanced training data tailored to obtaining vendor-style sodium MRI reconstruction, such that the reconstructed image achieves an appearance consistent with the vendor-reconstructed image without requiring any post-processing. Besides, we devise a differentiable concentration mapping module containing B1 inhomogeneity correction, automatic phantom detection, and linear transformation to convert the reconstructed image into a TSC map. Experiments on 107 subjects show that our method outperforms the existing representative methods by 7.19 dB in PSNR for image reconstruction and 2.54 mM in MAE for TSC quantification. Prospective evaluation on three additional subjects demonstrates the feasibility of SoReCon under real-world four-fold acceleration, reducing the primary sodium acquisition time from 15 min to 3.75 min and highlighting its potential for clinical translation. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/4553_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository N/A Link to the Dataset(s) N/A BibTex @InProceedings{CaoYil_A3D_MICCAI2026,         author = { Cao, Yilin AND Duan, Caohui AND Shen, Dinggang AND Lou, Xin AND Sun, Kaicong},         title = { { A 3D Unrolling Framework for Joint Sodium MRI Reconstruction and Concentration Quantification } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16888},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper The authors present a deep-learning based reconstruction method for accelerated sodium MRI. The networks enable end-to-end total sodium concentration (TSC) mapping from raw k-space data. In-vivo studies show that the proposed method preserves both reconstruction fidelity and quantitative accuracy at a fourfold acceleration factor. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The proposed method directly produces TSC images, which is of most interest in sodium MRI. In the current practice, deriving TSC from sodium MRI relies on simple (linear model) but tedious post-processing, often including manual masking for calibration phantoms. This work achieve automation for the post-processing step. 2.This work also implemented bilateral filter-based reconstruction for use as the reference. The reference seemingly suppresses noise while preserving edges. 3.Prospective study shows that the proposed method outperformed scanner’s reconstruction pipeline. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.p4: What is “clinical DICOM reconstruction”? Is it named after the image format? This is a very confusing nomenclature. 2.The results are unconvincing and the mechanisms of the proposed method is not fully explored. Please refer to the comments below. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html 1.The authors claim that non-uniform FFT (NUFFT) reconstruction produced blur images. However, NUFFT as a Fourier transform variant does not inherently cause blurring (c. f. Behl et. al. , 2016, Magnetic Resonance in Medicine for clear 23Na images by NUFFT). Instead, the blurring might be rooted from density compensation. Authors should describe how the density compensation weighting was performed during NUFFT. Additionally, the acronym should be defined at its first occurrence in the manuscript. 2.The authors should explain how the retrospective data set was divided into train/test sets. 3.The effect of extra B1 correction was not fully discussed. (1) Please introduce NAA/NAV with a bit more details. (2) How much time does the extra scan cost? 4.The CMM is not justified based on the presented evidence. The ablation study shows that the proposed CMM leads to lower PSNR and SSIM. Also, the authors should provide the performance in the case of 2x acceleration. 5.Fig 5: It would be interesting to compare the proposed method compared with the bilateral filtered reconstruction on the undersampled k-space data. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? While this work presents an interesting method, the experiment design is probably flawed and there lacks critical evidence supporting authors’ claim. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Reject [Post rebuttal] Please justify your final decision from above. The authors have provided responses to questions regarding the soundness of results. However, I am afraid extra work is required to understand and verify the proposed design, particularly: (1) the improvment seen after adding CMM and the effect of incorporating data from prolonged auxilliary scans; (2) seemingly incoherent ablation study results; (3) unexplained lower performance on prospective data when compared to the test set. Review #2 Please describe the contribution of the paper The main contributions include: 1) The first end-to-end framework from undersampled k-space to quantitative sodium maps by jointly performing sodium MRI reconstruction and sodium concentration quantification; 2) Construction of a k-space training set for sodium MRI reconstruction with enhanced sharpness; 3) Integrating a concentration mapping module and concentration loss into network training to enable concentration error backpropagation; 4) Evaluated on a relatively large in-house data containing 110 patients including a prospective study, reducing scan time from 15 minutes to 3.75 minutes (4×) while preserving reconstruction quality. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. This paper presents several notable strengths. First, it addresses a clinically important problem in sodium MRI by tackling both low SNR and long acquisition time through a unified framework. Second, the proposed end-to-end approach that jointly performs image reconstruction and concentration quantification from undersampled k-space is well-motivated and represents a meaningful departure from conventional two-stage pipelines. Third, the integration of a differentiable concentration mapping module incorporating B1 correction and phantom calibration enhances the physiological relevance of the method. In addition, the study is supported by a relatively large in-house dataset and includes prospective validation, which strengthens its practical significance. Finally, the demonstrated acceleration with preserved reconstruction fidelity highlights the potential for real clinical translation. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. The limited prospective validation restricts evidence of generalizability. Additionally, the ablation study does not fully justify the contribution of individual components. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (5) Accept — should be accepted, independent of rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The first end-to-end framework from undersampled k-space to quantitative sodium maps; Evaluated on a relatively large in-house data; A prospective study, demonstrating the reduction of scan time from 15 minutes to 3.75 minutes. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The first end-to-end framework from undersampled k-space to quantitative sodium maps; Evaluated on a relatively large in-house data; A prospective study, demonstrating the reduction of scan time from 15 minutes to 3.75 minutes. Review #3 Please describe the contribution of the paper The paper proposes a new learning technique for 23Na reconstruction. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. *The approach is new *Evaluations are performed with real prospective data Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. *The paper does not compare against reasonable baseline techniques. *The paper focuses on PSNR, SSIM, and MAE without identifying limitations. This approach has been recently criticized. (DOI:10.1002/mrm.70377) *The literature review has a major omission with respect to sodium reconstruction methods that use 1H reference images for denoising and edge preservation: Atkinson IC, Thulborn KR, Lu A, Haldar J, Zhou XJ, Claiborne T, Liang ZP. Quantitative 23-sodium and 17-oxygen MR imaging in human brain at 9.4 Tesla enhanced by constrained k-space reconstruction. In Proceedings of the 16th Annual Meeting of ISMRM, Toronto 2008 (p. 335). Gnahm, C., Bock, M., Bachert, P., Semmler, W., Behl, N.G.R. and Nagel, A.M. (2014), Iterative 3D projection reconstruction of 23Na data with an 1H MRI constraint. Magn. Reson. Med., 71: 1720-1732. Gnahm C, Nagel AM. Anatomically weighted second-order total variation reconstruction of 23Na MRI using prior information from 1H MRI. Neuroimage. 2015 Jan 15;105:452-61. Lachner S, Zaric O, Utzschneider M, Minarikova L, Zbýň Š, Hensel B, Trattnig S, Uder M, Nagel AM. Compressed sensing reconstruction of 7 Tesla 23Na multi-channel breast data using 1H MRI constraint. Magnetic resonance imaging. 2019 Jul 1;60:145-56. Zhao Y, Guo R, Li Y, Thulborn KR, Liang ZP. High-resolution sodium imaging using anatomical and sparsity constraints for denoising and recovery of novel features. Magnetic Resonance in Medicine. 2021 Aug;86(2):625-36. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The novelty is borderline, the results are hard to interpret, and the paper has cited almost none of the relevant 23Na reconstruction methods. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The authors responded well to comments Author Feedback We thank the meta-reviewer and reviewers for your constructive feedback and positive recognition of our work on end-to-end sodium MRI reconstruction and quantitative concentration mapping. Below we clarify the major concerns. (1) Novelty, related work, and baseline positioning (R1, R3, MR). We thank R3 for highlighting prior 1H-constrained methods (e.g., Gnahm et al.), which use external anatomical priors. Since our method relies on multi-coil k-space data for reconstruction rather than on T1w priors for postprocessing (denoising), our protocol did not acquire 1H data and we cannot use these 1H-constrained methods for comparison. Instead, we selected three baselines: NUFFT, compressed sensing, and 3D U-Net, covering both traditional and learning-based approaches and allowing fair comparison. We will add the 1H-constrained methods in the related work. We thank R1 for suggesting the bilateral-filtered reconstruction and will add it into comparison in our journal version. (2) Clarification of NUFFT/regridding and terminology (R1, MR). We apologize for the previous unclarity. For NUFFT, we used Pipe-Menon algorithm for density compensation and chose the best parameters (e.g., iteration number) for the highest PSNR/SSIM, yielding relatively blurred reconstruction. For training data preparation, we used regridding on raw data without density compensation and took it as input for our sharpening procedure, which balances noise suppression and structural visibility. We will release the code upon acceptance. We agree that “clinical DICOM reconstruction” is confusing and will replace it with “scanner-provided reconstruction”. We will clarify the above issues in revised version. (3) Justification of CMM and evaluation metrics (R1, R2, R3, MR). Our end-to-end framework aims to obtain both accurate sodium quantification and reconstruction fidelity. At 4× acceleration, CMM substantially improves TSC MAE (9.17→6.16 mM) with slight PSNR/SSIM changes, indicating 1) TSC is easier to recover than image reconstruction for high acceleration, and the model balances both; 2) reconstruction improvement does not always align with TSC [Haldar et al.,2026]. Both show the importance of end-to-end learning. At 2× acceleration, CMM improves both MAE (7.79→6.20 mM) and PSNR (39.55→40.16 dB). We did not show the results due to page limits. Besides PSNR/SSIM, we included TSC error, prospective validation, and visual comparison. Prospective images were reviewed by experienced radiologists and rated as excellent compared to scanner-provided images, showing its clinical potential. We will discuss metric limitations and refer to [Haldar et al.,2026] in the revised manuscript. (4) Experimental protocol and B1 correction (R1, MR). The retrospective dataset contains 110 cases. 3 were excluded due to quality control. The remaining 107 cases were split at subject level into 77/10/20 for training/validation/test. For B1 calibration, due to B1+ transmit field inhomogeneity, image intensity may not accurately reflect the underlying sodium concentration. Therefore, we acquired high-SNR NAA and homogeneous NAV acquisitions and the voxel-wise NAA/NAV ratio estimates the relative transmit sensitivity map for quantitative bias correction. Our protocol includes two NAA scans for improved SNR (1 min each) and one NAV scan (6 min). These scans are part of our standard workflow. We will add these details to the revised version. (5) Prospective validation and practical feasibility (R2, MR). We acknowledge that the prospective data is limited due to high cost and long acquisition time. Nevertheless, we performed quantitative and visual evaluation on these prospective acquisitions and will collect more data in future to strengthen the clinical utility. From our empirical results, we found end-to-end learning brings the largest gain and CMM the second. Due to page limits, this was omitted but will be added in the revised version. Meta-Review Meta-review #1 Your recommendation Invite for Rebuttal Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. There seems to be concerns around interpretability of the results, innovation, and literature review, resulting in mixed reviews. Strengths: The work addresses an important sodium MRI problem by proposing an end-to-end 3D unrolled framework that jointly reconstructs undersampled k-space and produces total sodium concentration maps, incorporating B1 correction, phantom calibration, a relatively large in-house dataset, and prospective validation. Weaknesses: Reviewers found the experimental evidence uneven, with limited prospective validation, incomplete ablation of components, unclear or possibly flawed experimental design, missing comparisons to strong or relevant sodium reconstruction baselines, and an insufficient literature review that makes the novelty hard to assess. The authors therefore be invited for the rebuttal. After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. The authors are advised to address the reviewers’ remaining concerns and do what they promise to do in their rebuttal. Meta-review #2 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. This paper proposes a 3D unrolling framework for joint sodium MRI reconstruction and concentration quantification, aiming to reconstruct accelerated sodium MRI from undersampled k-space while directly producing quantitative total sodium concentration maps. The reviewers recognized the clinical importance of accelerated sodium MRI, the value of an end-to-end framework from k-space to quantitative sodium maps, and the inclusion of B1 correction, phantom calibration, a relatively large in-house dataset, and prospective validation. They also appreciated the potential reduction of scan time from 15 minutes to 3.75 minutes. The reviewers raised concerns about unclear experimental design, limited prospective validation, incomplete ablation, missing sodium MRI baseline comparisons, and insufficient discussion of related 1H-constrained sodium reconstruction methods. In the rebuttal, the authors clarified the NUFFT/regridding setup, dataset split, B1 calibration protocol, the role of the concentration mapping module, and the rationale for selected baselines. They also committed to improving the related-work discussion and terminology. Although some limitations remain, especially the limited prospective data and the need for clearer ablation presentation, the rebuttal sufficiently addressed the main concerns, and most reviewers supported acceptance after rebuttal. Therefore, I recommend acceptance. Meta-review #3 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. Post-rebuttal consensus is 2 accept and 1 reject, with R3 moving from weak reject to weak accept. The paper is not a clear accept because R1’s technical objections remain meaningful, but the rebuttal improved the balance. I would recommend accept, emphasizing that the final version must include the promised clarifications, related work, metric limitations, and prospective validation caveats. Meta-review #4 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. This paper rebuilds undersampled sodium MRI and outputs tissue sodium concentration maps in one trained pass. The main worry at review was that the ablation made no sense, since the concentration module lowers PSNR and SSIM. The rebuttal clears this up. The panel’s reconstruction expert read the same point and moved from reject to accept. The rest holds up from the submitted paper. There are three real baselines, and the extra methods one reviewer asked for need 1H scans this study never collects. PSNR and SSIM are weak metrics, but they are not the headline. The real number is concentration error in mm, which is the right thing to measure and it improves. The work also has real clinical weight, with a large in-house dataset and a prospective study rated by radiologists. What is left is fixable at camera-ready: the 3.75 minute figure leaves out about 8 minutes of calibration scans and should be given as total time, the prospective set is weaker than the test set and needs a line of explanation, and the “first” claim is only true because it is limited to sodium, since the joint reconstruct-and-quantify idea itself comes from proton qMRI. It is a sound, clinically useful paper at borderline accept bar. back to top</summary></entry><entry><title type="html">A Clinical Guideline-Grounded Hybrid Agentic Framework for Holistic Epilepsy Management</title><link href="https://papers.miccai.org/miccai-2026/0004-Paper2852" rel="alternate" type="text/html" title="A Clinical Guideline-Grounded Hybrid Agentic Framework for Holistic Epilepsy Management" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0004-Paper2852</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0004-Paper2852">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;Epilepsy is a chronic neurological disorder requiring multi-faceted management, including seizure detection, syndrome diagnosis, prognostication, antiseizure medication recommendation, epileptogenic zone localization, and surgical outcome prediction. Although numerous deep learning approaches have been developed for individual tasks, these models are typically siloed and modality-specific (e.g., EEG for seizure detection, MRI for localization), failing to reflect the multidisciplinary nature of real-world epilepsy care, where epileptologists, neuroradiologists, neurosurgeons, neuropsychologists and neuropsychiatrists jointly interpret heterogeneous evidence to guide decisions. In this work, we propose a clinical guideline-grounded hybrid multi-agent framework for holistic epilepsy management. Heterogeneous patient data is processed through modality-specific discriminative and generative models, where textual interpretations from generative agents are combined with structured predictions from discriminative models as auxiliary guidance. This aggregated evidence is passed to a central orchestrating agent grounded in international epilepsy guidelines, which evaluates multi-modal findings within structured clinical pathways and performs iterative cross-agent coordination for evidence-informed decision-making. We evaluate our framework across two datasets spanning six epilepsy management tasks and also introduce a publicly available multi-modal, multi-task epilepsy benchmark. Results demonstrate that integrating discriminative evidence with guideline-grounded generative coordination yields more reliable and comprehensive decisions compared to conventional LLM-based and task-specific baselines. Our dataset and code is available at https://github.com/khoapham154/epi_guide.git.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/2852_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/2852_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/khoapham154/epi_guide&quot;&gt;https://github.com/khoapham154/epi_guide&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{PhaDuy_AClinical_MICCAI2026,
        author = { Pham, Duy Khoa AND Giritharan, Dinesh AND Camargo de Oliveira, Guilherme AND Vo, Bao Quoc AND Verspoor, Karin AND Law, Meng AND Kwan, Patrick AND Ge, Zongyuan AND Mehta, Deval},
        title = { { A Clinical Guideline-Grounded Hybrid Agentic Framework for Holistic Epilepsy Management } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16895},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper introduces EPI-GUIDE, the first clinical guideline-based hybrid multi-agent framework for the holistic epilepsy management. To address the limitations of existing epilepsy AI models—specifically their unimodal and single-task constraints—as well as the hallucination issues inherent in purely generative large models, this framework couples discriminative models (which provide structured predictions) with generative agents (which produce textual explanations). These elements are integrated by a guideline-driven orchestrating agent that synthesizes multimodal evidence to generate the final output. Furthermore, the authors constructed and evaluated a multimodal, multi-task epilepsy dataset involving 306 patients, with plans for its open-source release.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1、Existing medical large language models often suffer from weak numerical reasoning and a high propensity for hallucinations. This paper mitigates these issues by introducing discriminative models to provide probabilistic metrics and predictive results, which are then converted into textual evidence for the LLM. This design retains the high precision of deep learning in image/signal processing while leveraging the LLM’s logical reasoning, closely mimicking the multidisciplinary team consultation paradigm in real-world clinical practice. 
2、Before making a final decision, the system cross-references multimodal evidence with standardized clinical pathways retrieved via RAG. If contradictions arise, the system triggers a “FOLLOW-UP” mechanism for multi-turn dialogue to resolve conflicts. This fully transparent reasoning process not only enhances performance but also significantly bolsters clinician trust in the model’s outputs.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1、Limited depth in multimodal representation fusion. The current framework utilizes a typical “Late Fusion” approach, where each modality independently generates text or predictive probabilities before being concatenated and fed to the central orchestrating agent. While this facilitates LLM processing, it lacks deep cross-modal representation learning in the feature space. Establishing deep correlations between structural brain imaging and functional EEG signals within a latent space could potentially capture more intricate pathogenic mechanisms, further improving performance in complex neurological diagnoses. 
2、Insufficient robustness analysis for missing modalities. Although the authors note that the MME dataset reflects clinical diversity (i. e. , only 94 of 306 patients have MRI, and 71 have EEG), the paper lacks a detailed analysis of how the specific agents and the central coordinator dynamically adjust their prompts or confidence thresholds when a core modality is missing. 
3、Undefined conflict resolution mechanism. There is a potential for contradictory predictions between the generative and discriminative models. The consistency check performed by the coordinator currently lacks quantitative standards or a formalized arbitration logic. 
4、Dataset bias and incomplete metrics. The MME dataset exhibits significant distribution bias (data imbalance). However, the paper only reports accuracy, omitting critical balanced metrics such as F1-score and recall. Without these, it is impossible to verify the model’s effectiveness on minority classes.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission has provided an anonymized link to the source code, dataset, or any other dependencies.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The limited method novelty and unclear writing leads to the overall score of this paper.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents EPI-GUIDE, a new multi-agent system built around clinical guidelines to help manage epilepsy. It brings together two main parts: specialized models that handle specific tasks and give clear, quantitative results, and large language model agents that coordinate everything, making decisions based on international epilepsy guidelines. Together, they tackle a range of jobs-like figuring out seizure types, pinpointing where seizures start in the brain, recommending medications, and predicting how patients will do after surgery. Everything works as a unified system, drawing on different types of information to support decision-making for epilepsy care.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;++ combines the reliability of discriminative models with the flexibility of generative agents, addressing the instability and hallucination issues of pure LLMs 
++ the orchestrating agent uses established international epilepsy guidelines
++ Results show significant accuracy improvements across all six tasks&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;

      &lt;ul&gt;
        &lt;li&gt;integrating multiple discriminative models, LLM agents and a central orchestrator likely requires substantial hardware resources, which could hinder deployment in low‑resource clinical settings.&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission has provided an anonymized link to the source code, dataset, or any other dependencies.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;++ combines the reliability of discriminative models with the flexibility of generative agents, addressing the instability and hallucination issues of pure LLMs 
++ the orchestrating agent uses established international epilepsy guidelines
++ Results show significant accuracy improvements across all six tasks
But - integrating multiple discriminative models, LLM agents and a central orchestrator likely requires substantial hardware resources, which could hinder deployment in low‑resource clinical settings.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Somewhat confident (2)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes EPI-GUIDE, a hybrid multi-agent framework for holistic epilepsy management. The system integrates modality-specific discriminative models (ResNet-50, MedSigLIP, PubMedBERT, REVE) with generative LLM agents (MedGemma variants), coordinated by a RAG-grounded orchestrating agent (GPT-OSS-120b) that is anchored to international epilepsy clinical guidelines (ILAE, NICE).&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The framework addresses six epilepsy management tasks across two datasets — a curated multi-modal multi-task epilepsy (MME) cohort and a public HD-EEG dataset. Ablation studies confirm the contribution of both discriminative auxiliary guidance and RAG-based guideline grounding. Expert neurologist evaluation on 15 cases further validates clinical utility.
﻿&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Section 2.2 states that the orchestrator issues FOLLOW-UP queries “to resolve discrepancies” when evidence is inconsistent, and may perform up to three multi-turn rounds. However, the termination condition for transitioning from FOLLOW-UP to COMPLETE is not formally specified beyond the qualitative description in Equation (4). Could the authors provide the explicit decision rule by which the orchestrator determines that sufficient concordance has been reached, and report statistics on how frequently each number of turns (1, 2, 3) was required across different tasks and datasets?&lt;/p&gt;

      &lt;p&gt;2.The ablation study in Table 3 evaluates four configurations by removing discriminative guidance and/or RAG independently. However, two important ablation dimensions are absent: first, the contribution of the multi-turn interaction mechanism itself (i.e., replacing iterative FOLLOW-UP queries with a single-turn orchestration); and second, the effect of the specific guideline corpus (e.g., replacing ILAE/NICE guidelines with a general medical knowledge base or no grounding). Without these ablations, it is difficult to attribute the 7.0-point gain from RAG specifically to epilepsy-specific guideline content rather than to the retrieval mechanism in general.&lt;/p&gt;

      &lt;p&gt;3.The HD-EEG public dataset comprises only 7 subjects and 61 stimulation sessions, which is an extremely small sample for evaluating a multi-agent framework. The standard deviations reported in Table 2 are correspondingly large (e.g., EZ localization: 64.1±5.2 for EPI-GUIDE vs. 60.8±14.1 for REVE), suggesting high variance across folds.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;While this paper tackles an important problem and presents some promising results, there are several key issues that need to be addressed before it can be considered in its current form.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-4&quot;&gt;Review #4&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes EPI-GUIDE, a clinical guideline-grounded hybrid agentic framework for holistic epilepsy management. The framework combines modality-specific discriminative models and modality-specific generative agents, converts discriminative outputs into textual auxiliary evidence, and then uses a central orchestrating agent grounded in international epilepsy guidelines to integrate multimodal evidence and make task-specific decisions. The paper evaluates the framework across two datasets and six epilepsy-related tasks, including epilepsy type classification, seizure type classification, epileptogenic zone localization, antiseizure medication response prediction, surgical outcome prediction, and stimulation-related tasks. The work also introduces a curated multi-modal, multi-task epilepsy benchmark intended for public release.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The paper addresses a clinically important and genuinely multi-disciplinary problem. Epilepsy management is not a single-task prediction problem, and the manuscript correctly recognizes the need to integrate EEG, MRI, clinical data, and guideline-based reasoning.
2.The hybrid design is well motivated. The combination of discriminative models for robust modality-specific prediction and generative agents for coordination and explanation is sensible, especially given the known instability of purely generative multi-agent systems in medicine.
3.Guideline grounding is a meaningful contribution. Rather than relying on unconstrained LLM discussion, the framework explicitly uses epilepsy guidelines and textbook knowledge to structure the orchestrator’s decision pathway.
4.The paper evaluates the framework on multiple clinically relevant tasks rather than only one narrow benchmark. This broadens the practical relevance of the work and better reflects real-world epilepsy workflows.
5.The experimental results are promising. On the MME dataset, EPI-GUIDE improves over both strong discriminative baselines and generative baselines, and on the HD-EEG dataset it also outperforms the listed alternatives.
6.The ablation study is informative and supports the value of both discriminative auxiliary guidance and guideline-grounded retrieval.
7.The inclusion of qualitative expert review, although limited in scale, is a useful step toward clinical validation.
8.The manuscript is clearly organized and the framework diagrams are helpful.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The novelty should be positioned more carefully. Multi-agent and multi-modal medical reasoning frameworks have already been explored in prior work, including MDAgents [12], Agent Hospital [14], Pathfinder [6], and MedAgent-Pro [25]. In addition, epilepsy-specific foundation-model or LLM-based reasoning efforts such as epilepsy drug recommendation [22] and EZ localization from semiology using LLMs [28] have already started to appear. Therefore, the present work is better viewed as an important epilepsy-focused hybrid and guideline-grounded extension, rather than an entirely new starting point.
2.The claim of “holistic” epilepsy management should be interpreted with some caution. Although the framework spans multiple tasks, the actual data availability is still limited and incomplete: the curated cohort contains 306 patients, but MRI is available for only 94 and EEG for only 71, and several tasks have substantially fewer labeled cases. This makes the multimodal and multi-task setting clinically meaningful, but still relatively sparse in practice.
3.The fairness of the baseline comparison is not fully clear. EPI-GUIDE is a system-level framework combining multiple discriminative models, generative agents, guideline retrieval, and multi-turn orchestration, while several baselines are single-modality or single-model methods. The comparison is still useful, but it makes it difficult to isolate whether the gains come primarily from hybrid evidence integration, guideline grounding, or simply greater system complexity.
4.The paper would benefit from stronger detail on the benchmark construction and data split design. Because the MME dataset is heterogeneous and partially missing across tasks, more clarity is needed on how folds were generated, whether missing modalities create information imbalance across methods, and how stable the reported results are under this sparsity.
5.The expert evaluation is a positive addition but still limited in scale. Fifteen cases and 61 task-level judgments are encouraging as an initial qualitative validation, but this is not yet sufficient to strongly support broader claims about clinical reliability and utility.
6.Some claims about “more reliable and comprehensive decisions” are directionally plausible, but they currently rely mainly on task accuracy and a small-scale expert study. Additional error analysis, discordance analysis, or calibration-style evaluation would make these claims more convincing.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors claimed to release the source code and/or dataset upon acceptance of the submission.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;I appreciate the authors for addressing epilepsy management as a genuinely multi-modal and multi-step clinical workflow rather than reducing it to a single prediction task. This is an important and underexplored perspective. I also think the paper makes a meaningful attempt to move beyond purely generative medical agents by introducing discriminative auxiliary evidence and explicit guideline grounding.
The overall framework is thoughtfully designed, and I found the hybrid formulation clinically intuitive. In particular, the idea that modality-specific predictors should provide structured evidence, while a central agent integrates that evidence under guideline constraints, is sensible and relevant for real-world decision support.
My main suggestions are aimed at strengthening the paper rather than questioning its overall direction. First, I encourage the authors to moderate the “first” claim and more explicitly position the work relative to existing multi-agent medical reasoning and epilepsy-specific AI systems. Second, the paper would benefit from a more detailed account of the MME cohort, including missing-modality patterns and how these affect the folds and task-specific comparisons. Third, I would welcome stronger analysis of what exactly the orchestrator contributes beyond strong modality-specific models, especially in discordant or incomplete-information cases. Finally, the expert evaluation is a good start, but expanding the qualitative/clinical assessment would substantially strengthen the claims of reliability and practical value.
Overall, I found the paper promising, clinically relevant, and more mature than many purely agentic medical AI submissions.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;My recommendation is based on the paper’s clear clinical motivation, coherent system design, and promising empirical performance, balanced against some limitations in novelty positioning and validation depth.&lt;/p&gt;

      &lt;p&gt;I think this work addresses an important gap in current medical AI for epilepsy. Rather than treating each task in isolation, the paper proposes a hybrid framework that integrates modality-specific discriminative evidence, generative interpretation, and explicit guideline-grounded orchestration. This is a meaningful direction, and the results across two datasets and multiple tasks are encouraging. The ablation study also supports the importance of both discriminative auxiliary evidence and guideline grounding.&lt;/p&gt;

      &lt;p&gt;At the same time, I think some claims should be framed more carefully. The novelty is partly incremental relative to prior multi-agent and multimodal medical reasoning work, the curated dataset is still relatively small and incomplete across modalities/tasks, and the expert evaluation is limited in scale. I would also have liked a clearer decomposition of what the orchestrating agent contributes beyond strong task-specific predictors.&lt;/p&gt;

      &lt;p&gt;Despite these limitations, I believe the paper makes a useful and timely contribution. It is clinically grounded, experimentally broader than many competing submissions, and points toward an important direction for hybrid agentic medical decision support. For these reasons, I place it slightly above the acceptance threshold.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;After reading the authors’ rebuttal, I maintain my recommendation in favor of acceptance.
The rebuttal addresses several of my main concerns in a meaningful way. In particular, the authors clarified that they will soften the novelty claim and position the work more appropriately as a hybrid, guideline-grounded extension for epilepsy care rather than as a wholly new paradigm. They also substantially improved the transparency of the orchestration mechanism by specifying the confidence-weighted concordance rule, the COMPLETE/FOLLOW-UP threshold, the three-tier arbitration logic, and the distribution of required follow-up rounds on the MME dataset. In addition, they provided useful missing-modality analyses (including T-only, T+M, T+E, M-only, and E-only settings), clarified how missing modalities are handled through prompt branching and weight renormalization, and added more detail on the 5-fold split strategy and the omitted weighted/macro F1 results.
These clarifications strengthen my confidence in the paper. In particular, they make the hybrid system more interpretable at the decision level and reduce my earlier concern that some important implementation and evaluation details were underspecified. I also appreciate that the authors responded constructively by explicitly narrowing some of their claims, especially around novelty and the current scale of the expert reader study.
That said, some limitations remain. The expert evaluation is still relatively small, the HD-EEG benchmark remains limited in size, and the deeper contribution of multi-turn orchestration appears concentrated in harder cases rather than yielding a large overall gain. I also still view the current fusion strategy as relatively shallow compared with more deeply integrated multimodal approaches. However, these are limitations rather than fatal flaws, and in my view they do not outweigh the paper’s strengths: clear clinical motivation, a coherent hybrid design, meaningful use of guideline grounding, and broad evaluation across multiple epilepsy management tasks.
For these reasons, I support acceptance and keep the paper slightly above the acceptance threshold.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;Thank you for recognizing the importance of our work and your constructive feedback. If we have addressed your concerns, we kindly ask to reconsider your score.&lt;/p&gt;

  &lt;p&gt;Open-source: EPIGUIDE (code, prompts, configs) and MME, to our knowledge the first multi-modal multi-task epilepsy benchmark, are released at the anonymous link.&lt;/p&gt;

  &lt;p&gt;1.Novelty, reader study (R#4): Our novelty lies in the hybrid discriminative-generative coupling, guideline grounding via expert-curated ILAE documents, and multi-task epilepsy management. We will position this as a hybrid, guideline-grounded epilepsy care extension rather than an entirely new paradigm in the revision. We acknowledge the limited scale of our expert reader study, but the large gap (93.4% vs 54.1% task accuracy; 3.64 vs 2.64 Likert for EPIGUIDE vs LLM-only) suggest “promising preliminary results in assistive epilepsy management”. We will revise claims accordingly and outline larger scale studies, error analysis, and reasoning evaluation as future work.&lt;/p&gt;

  &lt;p&gt;2.Orchestrator conflict resolution (R#1,R#3,R#4): For each patient/task, the orchestrator compares the classifier’s top-1 prediction (with confidence) to modality agents’ top-1 labels, aggregating agreement into a confidence-weighted patient-level concordance score. If ≥0.7 (empirically selected threshold), the case is marked COMPLETE; otherwise, a FOLLOW-UP targets the most disagreed/lowest-confidence task (≤3 rounds; 44.4%|42.5%|13.1% cases required 1|2|3 rounds on MME). Final predictions from orchestrator apply a three-tier rule: trust the classifier if confidence ≥85%, defer to the generative model if &amp;lt;50%, else apply a confidence-weighted vote. On the 22–45% disagreement cases, the three-tier rule yields a mF1 of 65% vs mF1 of 27% for generative-only orchestration, demonstrating the value of hybrid approach.&lt;/p&gt;

  &lt;p&gt;3.Missing modality, orchestrator ablations (R#1,R#3,R#4): On all 306 patients, EPIGUIDE under reduced modalities (were omitted due to brevity) achieved T-only 80.1%, T+M 83.4%, T+E 82.6%; M-only (29.9% on the 94-MRIs) and E-only (50.6% on the 71 EEGs) confirm text as the anchor modality, with imaging adding complementary gains. The orchestrator handles missing modality via a “No-data” prompt branch, renormalised modality weights, and confidence-tier routing to the generative model. Replacing FOLLOW-UP with single-pass yields negligible overall gain (+0.1%), but +6.3% on cases needing ≥2 rounds, indicating benefits for harder cases. The +7.0% RAG gain decomposes into +3.0% generic retrieval of topic-matched Pubmed articles and +4.0% ILAE/NICE-specific content, with the epilepsy-specific gains concentrated on EZ localization and surgical outcome.&lt;/p&gt;

  &lt;p&gt;4.MME split, metrics, variance (R#4,R#1,R#3): On 306 patients (5 tasks, missing modalities), our 5-fold CV is stratified by epilepsy type and modality availability (has_mri × has_eeg), ensuring balanced folds with no patient overlap. We computed wF1, mF1, but omitted for brevity. EPI-GUIDE achieves 84–86 wF1|78–80 mF1 vs 78.0|71.3 (best ensemble) and 56.8|43.1 (MedGemma-27B zero-shot), to be added to Table 1.The 6.6 wF1–mF1 gap is driven by low recall of &amp;lt;45% on minority classes: Multifocal EZ and On-treatment AED response. On HD-EEG, variance is higher due to the smaller cohort, though EPIGUIDE still shows lower variance than REVE (Table 2).&lt;/p&gt;

  &lt;p&gt;5.Simple fusion (R#1): Our choice deliberately mirrors real epilepsy MDT meetings where specialists interpret evidence independently before joint deliberation, and also to isolate our core novelty (see point 1). Table 3 confirms gains stem from discriminative evidence (+22.9) and guidelines (+7.0), not fusion depth. Missing modalities in the MME (MRI: 94/306, EEG: 71/306) complicate latent fusion, but we agree this is a promising future direction.&lt;/p&gt;

  &lt;p&gt;6.Compute (R#2): 8×A100 setup is for running GPT-OSS-120B locally; discriminative models are lightweight. Our modular design allows swapping in more efficient models, which we will note in the revision.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Invite for Rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes EPI-GUIDE, a guideline-grounded hybrid multi-agent framework for holistic epilepsy management, combining discriminative and generative agents with a central orchestrator. Reviewers recognize the clinical relevance, coherent design, and promising performance across multiple datasets and tasks.&lt;/p&gt;

      &lt;p&gt;The authors should focus on clarifying and addressing the following points: position the novelty more carefully relative to prior multi-agent and multimodal epilepsy AI systems; provide more details on the MME dataset, including missing-modality patterns, data splits, and their effect on task-specific comparisons; clarify the contribution of the orchestrator beyond strong modality-specific predictors, particularly in discordant or incomplete-information cases; expand the qualitative and clinical expert evaluation to strengthen claims of reliability and practical value; and frame claims about “holistic” management and novelty more cautiously given dataset limitations.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Reject&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper has clear clinical motivation and promising empirical results, but the reviews remain mixed and mainly weak. Concerns about limited methodological novelty, unclear writing, deployment complexity, and insufficient validation depth were not convincingly overcome. On balance, the work seems promising but not yet strong enough for acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-2&quot;&gt;Meta-review #2&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The review process has navigated a set of careful reviews that were tied between ‘weak reject’ and ‘weak accept’. But after consideration of the rebuttal, there has been a slight swing in the positive direction. Authors are encouraged to follow the careful reviews and constructive commentary.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-3&quot;&gt;Meta-review #3&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;In the rebuttal the authors addressed previously underspecified points such as specifying the exact novelty.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Pham, Duy Khoa AND Giritharan, Dinesh AND Camargo de Oliveira, Guilherme AND Vo, Bao Quoc AND Verspoor, Karin AND Law, Meng AND Kwan, Patrick AND Ge, Zongyuan AND Mehta, Deval</name></author><category term="Body -&gt; Brain" /><category term="Modalities -&gt; EEG / MEG / ECG / Physiological Signals" /><category term="Modalities -&gt; MRI" /><category term="Applications -&gt; Computer-Aided Diagnosis" /><category term="Applications -&gt; Multimodal Integration with Clinical / Genomic / Biomarkers" /><category term="Applications -&gt; Outcome Prediction / Prognosis / Longitudinal Modeling" /><category term="Machine Learning -&gt; Foundation Models" /><category term="Machine Learning -&gt; Multimodal Models / LLMs / VLMs" /><category term="Pham, Duy Khoa" /><category term="Giritharan, Dinesh" /><category term="Camargo de Oliveira, Guilherme" /><category term="Vo, Bao Quoc" /><category term="Verspoor, Karin" /><category term="Law, Meng" /><category term="Kwan, Patrick" /><category term="Ge, Zongyuan" /><category term="Mehta, Deval" /><summary type="html">Abstract Epilepsy is a chronic neurological disorder requiring multi-faceted management, including seizure detection, syndrome diagnosis, prognostication, antiseizure medication recommendation, epileptogenic zone localization, and surgical outcome prediction. Although numerous deep learning approaches have been developed for individual tasks, these models are typically siloed and modality-specific (e.g., EEG for seizure detection, MRI for localization), failing to reflect the multidisciplinary nature of real-world epilepsy care, where epileptologists, neuroradiologists, neurosurgeons, neuropsychologists and neuropsychiatrists jointly interpret heterogeneous evidence to guide decisions. In this work, we propose a clinical guideline-grounded hybrid multi-agent framework for holistic epilepsy management. Heterogeneous patient data is processed through modality-specific discriminative and generative models, where textual interpretations from generative agents are combined with structured predictions from discriminative models as auxiliary guidance. This aggregated evidence is passed to a central orchestrating agent grounded in international epilepsy guidelines, which evaluates multi-modal findings within structured clinical pathways and performs iterative cross-agent coordination for evidence-informed decision-making. We evaluate our framework across two datasets spanning six epilepsy management tasks and also introduce a publicly available multi-modal, multi-task epilepsy benchmark. Results demonstrate that integrating discriminative evidence with guideline-grounded generative coordination yields more reliable and comprehensive decisions compared to conventional LLM-based and task-specific baselines. Our dataset and code is available at https://github.com/khoapham154/epi_guide.git. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/2852_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository https://github.com/khoapham154/epi_guide Link to the Dataset(s) N/A BibTex @InProceedings{PhaDuy_AClinical_MICCAI2026,         author = { Pham, Duy Khoa AND Giritharan, Dinesh AND Camargo de Oliveira, Guilherme AND Vo, Bao Quoc AND Verspoor, Karin AND Law, Meng AND Kwan, Patrick AND Ge, Zongyuan AND Mehta, Deval},         title = { { A Clinical Guideline-Grounded Hybrid Agentic Framework for Holistic Epilepsy Management } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16895},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper This paper introduces EPI-GUIDE, the first clinical guideline-based hybrid multi-agent framework for the holistic epilepsy management. To address the limitations of existing epilepsy AI models—specifically their unimodal and single-task constraints—as well as the hallucination issues inherent in purely generative large models, this framework couples discriminative models (which provide structured predictions) with generative agents (which produce textual explanations). These elements are integrated by a guideline-driven orchestrating agent that synthesizes multimodal evidence to generate the final output. Furthermore, the authors constructed and evaluated a multimodal, multi-task epilepsy dataset involving 306 patients, with plans for its open-source release. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1、Existing medical large language models often suffer from weak numerical reasoning and a high propensity for hallucinations. This paper mitigates these issues by introducing discriminative models to provide probabilistic metrics and predictive results, which are then converted into textual evidence for the LLM. This design retains the high precision of deep learning in image/signal processing while leveraging the LLM’s logical reasoning, closely mimicking the multidisciplinary team consultation paradigm in real-world clinical practice. 2、Before making a final decision, the system cross-references multimodal evidence with standardized clinical pathways retrieved via RAG. If contradictions arise, the system triggers a “FOLLOW-UP” mechanism for multi-turn dialogue to resolve conflicts. This fully transparent reasoning process not only enhances performance but also significantly bolsters clinician trust in the model’s outputs. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1、Limited depth in multimodal representation fusion. The current framework utilizes a typical “Late Fusion” approach, where each modality independently generates text or predictive probabilities before being concatenated and fed to the central orchestrating agent. While this facilitates LLM processing, it lacks deep cross-modal representation learning in the feature space. Establishing deep correlations between structural brain imaging and functional EEG signals within a latent space could potentially capture more intricate pathogenic mechanisms, further improving performance in complex neurological diagnoses. 2、Insufficient robustness analysis for missing modalities. Although the authors note that the MME dataset reflects clinical diversity (i. e. , only 94 of 306 patients have MRI, and 71 have EEG), the paper lacks a detailed analysis of how the specific agents and the central coordinator dynamically adjust their prompts or confidence thresholds when a core modality is missing. 3、Undefined conflict resolution mechanism. There is a potential for contradictory predictions between the generative and discriminative models. The consistency check performed by the coordinator currently lacks quantitative standards or a formalized arbitration logic. 4、Dataset bias and incomplete metrics. The MME dataset exhibits significant distribution bias (data imbalance). However, the paper only reports accuracy, omitting critical balanced metrics such as F1-score and recall. Without these, it is impossible to verify the model’s effectiveness on minority classes. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission has provided an anonymized link to the source code, dataset, or any other dependencies. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The limited method novelty and unclear writing leads to the overall score of this paper. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #2 Please describe the contribution of the paper This paper presents EPI-GUIDE, a new multi-agent system built around clinical guidelines to help manage epilepsy. It brings together two main parts: specialized models that handle specific tasks and give clear, quantitative results, and large language model agents that coordinate everything, making decisions based on international epilepsy guidelines. Together, they tackle a range of jobs-like figuring out seizure types, pinpointing where seizures start in the brain, recommending medications, and predicting how patients will do after surgery. Everything works as a unified system, drawing on different types of information to support decision-making for epilepsy care. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. ++ combines the reliability of discriminative models with the flexibility of generative agents, addressing the instability and hallucination issues of pure LLMs  ++ the orchestrating agent uses established international epilepsy guidelines ++ Results show significant accuracy improvements across all six tasks Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. integrating multiple discriminative models, LLM agents and a central orchestrator likely requires substantial hardware resources, which could hinder deployment in low‑resource clinical settings. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission has provided an anonymized link to the source code, dataset, or any other dependencies. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? ++ combines the reliability of discriminative models with the flexibility of generative agents, addressing the instability and hallucination issues of pure LLMs  ++ the orchestrating agent uses established international epilepsy guidelines ++ Results show significant accuracy improvements across all six tasks But - integrating multiple discriminative models, LLM agents and a central orchestrator likely requires substantial hardware resources, which could hinder deployment in low‑resource clinical settings. Reviewer confidence Somewhat confident (2) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #3 Please describe the contribution of the paper This paper proposes EPI-GUIDE, a hybrid multi-agent framework for holistic epilepsy management. The system integrates modality-specific discriminative models (ResNet-50, MedSigLIP, PubMedBERT, REVE) with generative LLM agents (MedGemma variants), coordinated by a RAG-grounded orchestrating agent (GPT-OSS-120b) that is anchored to international epilepsy clinical guidelines (ILAE, NICE). Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The framework addresses six epilepsy management tasks across two datasets — a curated multi-modal multi-task epilepsy (MME) cohort and a public HD-EEG dataset. Ablation studies confirm the contribution of both discriminative auxiliary guidance and RAG-based guideline grounding. Expert neurologist evaluation on 15 cases further validates clinical utility. ﻿ Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.Section 2.2 states that the orchestrator issues FOLLOW-UP queries “to resolve discrepancies” when evidence is inconsistent, and may perform up to three multi-turn rounds. However, the termination condition for transitioning from FOLLOW-UP to COMPLETE is not formally specified beyond the qualitative description in Equation (4). Could the authors provide the explicit decision rule by which the orchestrator determines that sufficient concordance has been reached, and report statistics on how frequently each number of turns (1, 2, 3) was required across different tasks and datasets? 2.The ablation study in Table 3 evaluates four configurations by removing discriminative guidance and/or RAG independently. However, two important ablation dimensions are absent: first, the contribution of the multi-turn interaction mechanism itself (i.e., replacing iterative FOLLOW-UP queries with a single-turn orchestration); and second, the effect of the specific guideline corpus (e.g., replacing ILAE/NICE guidelines with a general medical knowledge base or no grounding). Without these ablations, it is difficult to attribute the 7.0-point gain from RAG specifically to epilepsy-specific guideline content rather than to the retrieval mechanism in general. 3.The HD-EEG public dataset comprises only 7 subjects and 61 stimulation sessions, which is an extremely small sample for evaluating a multi-agent framework. The standard deviations reported in Table 2 are correspondingly large (e.g., EZ localization: 64.1±5.2 for EPI-GUIDE vs. 60.8±14.1 for REVE), suggesting high variance across folds. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? While this paper tackles an important problem and presents some promising results, there are several key issues that need to be addressed before it can be considered in its current form. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #4 Please describe the contribution of the paper This paper proposes EPI-GUIDE, a clinical guideline-grounded hybrid agentic framework for holistic epilepsy management. The framework combines modality-specific discriminative models and modality-specific generative agents, converts discriminative outputs into textual auxiliary evidence, and then uses a central orchestrating agent grounded in international epilepsy guidelines to integrate multimodal evidence and make task-specific decisions. The paper evaluates the framework across two datasets and six epilepsy-related tasks, including epilepsy type classification, seizure type classification, epileptogenic zone localization, antiseizure medication response prediction, surgical outcome prediction, and stimulation-related tasks. The work also introduces a curated multi-modal, multi-task epilepsy benchmark intended for public release. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The paper addresses a clinically important and genuinely multi-disciplinary problem. Epilepsy management is not a single-task prediction problem, and the manuscript correctly recognizes the need to integrate EEG, MRI, clinical data, and guideline-based reasoning. 2.The hybrid design is well motivated. The combination of discriminative models for robust modality-specific prediction and generative agents for coordination and explanation is sensible, especially given the known instability of purely generative multi-agent systems in medicine. 3.Guideline grounding is a meaningful contribution. Rather than relying on unconstrained LLM discussion, the framework explicitly uses epilepsy guidelines and textbook knowledge to structure the orchestrator’s decision pathway. 4.The paper evaluates the framework on multiple clinically relevant tasks rather than only one narrow benchmark. This broadens the practical relevance of the work and better reflects real-world epilepsy workflows. 5.The experimental results are promising. On the MME dataset, EPI-GUIDE improves over both strong discriminative baselines and generative baselines, and on the HD-EEG dataset it also outperforms the listed alternatives. 6.The ablation study is informative and supports the value of both discriminative auxiliary guidance and guideline-grounded retrieval. 7.The inclusion of qualitative expert review, although limited in scale, is a useful step toward clinical validation. 8.The manuscript is clearly organized and the framework diagrams are helpful. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.The novelty should be positioned more carefully. Multi-agent and multi-modal medical reasoning frameworks have already been explored in prior work, including MDAgents [12], Agent Hospital [14], Pathfinder [6], and MedAgent-Pro [25]. In addition, epilepsy-specific foundation-model or LLM-based reasoning efforts such as epilepsy drug recommendation [22] and EZ localization from semiology using LLMs [28] have already started to appear. Therefore, the present work is better viewed as an important epilepsy-focused hybrid and guideline-grounded extension, rather than an entirely new starting point. 2.The claim of “holistic” epilepsy management should be interpreted with some caution. Although the framework spans multiple tasks, the actual data availability is still limited and incomplete: the curated cohort contains 306 patients, but MRI is available for only 94 and EEG for only 71, and several tasks have substantially fewer labeled cases. This makes the multimodal and multi-task setting clinically meaningful, but still relatively sparse in practice. 3.The fairness of the baseline comparison is not fully clear. EPI-GUIDE is a system-level framework combining multiple discriminative models, generative agents, guideline retrieval, and multi-turn orchestration, while several baselines are single-modality or single-model methods. The comparison is still useful, but it makes it difficult to isolate whether the gains come primarily from hybrid evidence integration, guideline grounding, or simply greater system complexity. 4.The paper would benefit from stronger detail on the benchmark construction and data split design. Because the MME dataset is heterogeneous and partially missing across tasks, more clarity is needed on how folds were generated, whether missing modalities create information imbalance across methods, and how stable the reported results are under this sparsity. 5.The expert evaluation is a positive addition but still limited in scale. Fifteen cases and 61 task-level judgments are encouraging as an initial qualitative validation, but this is not yet sufficient to strongly support broader claims about clinical reliability and utility. 6.Some claims about “more reliable and comprehensive decisions” are directionally plausible, but they currently rely mainly on task accuracy and a small-scale expert study. Additional error analysis, discordance analysis, or calibration-style evaluation would make these claims more convincing. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The authors claimed to release the source code and/or dataset upon acceptance of the submission. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html I appreciate the authors for addressing epilepsy management as a genuinely multi-modal and multi-step clinical workflow rather than reducing it to a single prediction task. This is an important and underexplored perspective. I also think the paper makes a meaningful attempt to move beyond purely generative medical agents by introducing discriminative auxiliary evidence and explicit guideline grounding. The overall framework is thoughtfully designed, and I found the hybrid formulation clinically intuitive. In particular, the idea that modality-specific predictors should provide structured evidence, while a central agent integrates that evidence under guideline constraints, is sensible and relevant for real-world decision support. My main suggestions are aimed at strengthening the paper rather than questioning its overall direction. First, I encourage the authors to moderate the “first” claim and more explicitly position the work relative to existing multi-agent medical reasoning and epilepsy-specific AI systems. Second, the paper would benefit from a more detailed account of the MME cohort, including missing-modality patterns and how these affect the folds and task-specific comparisons. Third, I would welcome stronger analysis of what exactly the orchestrator contributes beyond strong modality-specific models, especially in discordant or incomplete-information cases. Finally, the expert evaluation is a good start, but expanding the qualitative/clinical assessment would substantially strengthen the claims of reliability and practical value. Overall, I found the paper promising, clinically relevant, and more mature than many purely agentic medical AI submissions. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? My recommendation is based on the paper’s clear clinical motivation, coherent system design, and promising empirical performance, balanced against some limitations in novelty positioning and validation depth. I think this work addresses an important gap in current medical AI for epilepsy. Rather than treating each task in isolation, the paper proposes a hybrid framework that integrates modality-specific discriminative evidence, generative interpretation, and explicit guideline-grounded orchestration. This is a meaningful direction, and the results across two datasets and multiple tasks are encouraging. The ablation study also supports the importance of both discriminative auxiliary evidence and guideline grounding. At the same time, I think some claims should be framed more carefully. The novelty is partly incremental relative to prior multi-agent and multimodal medical reasoning work, the curated dataset is still relatively small and incomplete across modalities/tasks, and the expert evaluation is limited in scale. I would also have liked a clearer decomposition of what the orchestrating agent contributes beyond strong task-specific predictors. Despite these limitations, I believe the paper makes a useful and timely contribution. It is clinically grounded, experimentally broader than many competing submissions, and points toward an important direction for hybrid agentic medical decision support. For these reasons, I place it slightly above the acceptance threshold. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. After reading the authors’ rebuttal, I maintain my recommendation in favor of acceptance. The rebuttal addresses several of my main concerns in a meaningful way. In particular, the authors clarified that they will soften the novelty claim and position the work more appropriately as a hybrid, guideline-grounded extension for epilepsy care rather than as a wholly new paradigm. They also substantially improved the transparency of the orchestration mechanism by specifying the confidence-weighted concordance rule, the COMPLETE/FOLLOW-UP threshold, the three-tier arbitration logic, and the distribution of required follow-up rounds on the MME dataset. In addition, they provided useful missing-modality analyses (including T-only, T+M, T+E, M-only, and E-only settings), clarified how missing modalities are handled through prompt branching and weight renormalization, and added more detail on the 5-fold split strategy and the omitted weighted/macro F1 results. These clarifications strengthen my confidence in the paper. In particular, they make the hybrid system more interpretable at the decision level and reduce my earlier concern that some important implementation and evaluation details were underspecified. I also appreciate that the authors responded constructively by explicitly narrowing some of their claims, especially around novelty and the current scale of the expert reader study. That said, some limitations remain. The expert evaluation is still relatively small, the HD-EEG benchmark remains limited in size, and the deeper contribution of multi-turn orchestration appears concentrated in harder cases rather than yielding a large overall gain. I also still view the current fusion strategy as relatively shallow compared with more deeply integrated multimodal approaches. However, these are limitations rather than fatal flaws, and in my view they do not outweigh the paper’s strengths: clear clinical motivation, a coherent hybrid design, meaningful use of guideline grounding, and broad evaluation across multiple epilepsy management tasks. For these reasons, I support acceptance and keep the paper slightly above the acceptance threshold. Author Feedback Thank you for recognizing the importance of our work and your constructive feedback. If we have addressed your concerns, we kindly ask to reconsider your score. Open-source: EPIGUIDE (code, prompts, configs) and MME, to our knowledge the first multi-modal multi-task epilepsy benchmark, are released at the anonymous link. 1.Novelty, reader study (R#4): Our novelty lies in the hybrid discriminative-generative coupling, guideline grounding via expert-curated ILAE documents, and multi-task epilepsy management. We will position this as a hybrid, guideline-grounded epilepsy care extension rather than an entirely new paradigm in the revision. We acknowledge the limited scale of our expert reader study, but the large gap (93.4% vs 54.1% task accuracy; 3.64 vs 2.64 Likert for EPIGUIDE vs LLM-only) suggest “promising preliminary results in assistive epilepsy management”. We will revise claims accordingly and outline larger scale studies, error analysis, and reasoning evaluation as future work. 2.Orchestrator conflict resolution (R#1,R#3,R#4): For each patient/task, the orchestrator compares the classifier’s top-1 prediction (with confidence) to modality agents’ top-1 labels, aggregating agreement into a confidence-weighted patient-level concordance score. If ≥0.7 (empirically selected threshold), the case is marked COMPLETE; otherwise, a FOLLOW-UP targets the most disagreed/lowest-confidence task (≤3 rounds; 44.4%|42.5%|13.1% cases required 1|2|3 rounds on MME). Final predictions from orchestrator apply a three-tier rule: trust the classifier if confidence ≥85%, defer to the generative model if &amp;lt;50%, else apply a confidence-weighted vote. On the 22–45% disagreement cases, the three-tier rule yields a mF1 of 65% vs mF1 of 27% for generative-only orchestration, demonstrating the value of hybrid approach. 3.Missing modality, orchestrator ablations (R#1,R#3,R#4): On all 306 patients, EPIGUIDE under reduced modalities (were omitted due to brevity) achieved T-only 80.1%, T+M 83.4%, T+E 82.6%; M-only (29.9% on the 94-MRIs) and E-only (50.6% on the 71 EEGs) confirm text as the anchor modality, with imaging adding complementary gains. The orchestrator handles missing modality via a “No-data” prompt branch, renormalised modality weights, and confidence-tier routing to the generative model. Replacing FOLLOW-UP with single-pass yields negligible overall gain (+0.1%), but +6.3% on cases needing ≥2 rounds, indicating benefits for harder cases. The +7.0% RAG gain decomposes into +3.0% generic retrieval of topic-matched Pubmed articles and +4.0% ILAE/NICE-specific content, with the epilepsy-specific gains concentrated on EZ localization and surgical outcome. 4.MME split, metrics, variance (R#4,R#1,R#3): On 306 patients (5 tasks, missing modalities), our 5-fold CV is stratified by epilepsy type and modality availability (has_mri × has_eeg), ensuring balanced folds with no patient overlap. We computed wF1, mF1, but omitted for brevity. EPI-GUIDE achieves 84–86 wF1|78–80 mF1 vs 78.0|71.3 (best ensemble) and 56.8|43.1 (MedGemma-27B zero-shot), to be added to Table 1.The 6.6 wF1–mF1 gap is driven by low recall of &amp;lt;45% on minority classes: Multifocal EZ and On-treatment AED response. On HD-EEG, variance is higher due to the smaller cohort, though EPIGUIDE still shows lower variance than REVE (Table 2). 5.Simple fusion (R#1): Our choice deliberately mirrors real epilepsy MDT meetings where specialists interpret evidence independently before joint deliberation, and also to isolate our core novelty (see point 1). Table 3 confirms gains stem from discriminative evidence (+22.9) and guidelines (+7.0), not fusion depth. Missing modalities in the MME (MRI: 94/306, EEG: 71/306) complicate latent fusion, but we agree this is a promising future direction. 6.Compute (R#2): 8×A100 setup is for running GPT-OSS-120B locally; discriminative models are lightweight. Our modular design allows swapping in more efficient models, which we will note in the revision. Meta-Review Meta-review #1 Your recommendation Invite for Rebuttal Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. This paper proposes EPI-GUIDE, a guideline-grounded hybrid multi-agent framework for holistic epilepsy management, combining discriminative and generative agents with a central orchestrator. Reviewers recognize the clinical relevance, coherent design, and promising performance across multiple datasets and tasks. The authors should focus on clarifying and addressing the following points: position the novelty more carefully relative to prior multi-agent and multimodal epilepsy AI systems; provide more details on the MME dataset, including missing-modality patterns, data splits, and their effect on task-specific comparisons; clarify the contribution of the orchestrator beyond strong modality-specific predictors, particularly in discordant or incomplete-information cases; expand the qualitative and clinical expert evaluation to strengthen claims of reliability and practical value; and frame claims about “holistic” management and novelty more cautiously given dataset limitations. After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Reject Please justify your recommendation. The paper has clear clinical motivation and promising empirical results, but the reviews remain mixed and mainly weak. Concerns about limited methodological novelty, unclear writing, deployment complexity, and insufficient validation depth were not convincingly overcome. On balance, the work seems promising but not yet strong enough for acceptance. Meta-review #2 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. The review process has navigated a set of careful reviews that were tied between ‘weak reject’ and ‘weak accept’. But after consideration of the rebuttal, there has been a slight swing in the positive direction. Authors are encouraged to follow the careful reviews and constructive commentary. Meta-review #3 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. In the rebuttal the authors addressed previously underspecified points such as specifying the exact novelty. back to top</summary></entry><entry><title type="html">A Confounder-aware Representation Learning Framework for Alzheimer’s Disease Classification</title><link href="https://papers.miccai.org/miccai-2026/0005-Paper1600" rel="alternate" type="text/html" title="A Confounder-aware Representation Learning Framework for Alzheimer’s Disease Classification" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0005-Paper1600</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0005-Paper1600">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;Alzheimer’s disease (AD) is the most prevalent neurodegenerative disorder, and current treatments cannot reverse its progression, making early diagnosis and intervention essential. Deep learning methods have shown promise for diagnosing AD from neuroimaging data, but these models are often affected by both visible and hidden sources of noise. Visible noise typically reflects systematic imaging differences, such as scanner-related biases and acquisition artifacts, whereas hidden noise is linked to individual variation that can lead models to depend on spurious patterns that do not generalize well. Most existing approaches emphasize statistical associations and do not adequately account for these confounding influences. Motivated by these limitations, we propose a multimodal, confounder-aware framework for AD diagnosis inspired by causal reasoning. The model includes a region-masking component (Mask Decoupling Module) that probes the image by making small, targeted changes to localized areas, then highlights the regions that most strongly affect the prediction, encouraging the network to prioritize structurally informative features. We also introduce a covariate-guided feature modulation mechanism that uses clinical and demographic variables to adjust local image features, thereby reducing variation attributable to these factors. Although the framework does not estimate causal effects directly, it is designed to learn more stable representations by limiting reliance on shortcuts, improving robustness and generalization in AD diagnostic models.
Our code is available at \url{https://anonymous.4open.science/r/Causal_fusion-4E6E}.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/1600_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/1600_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/redtea-code/Causal_fusion&quot;&gt;https://github.com/redtea-code/Causal_fusion&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{CheYua_AConfounderaware_MICCAI2026,
        author = { Chen, Yuanhao AND Chen, Jingwen AND Xu, Jinyi AND Yu, Haoqi AND Wang, Bin AND Li, Chunzhong AND Zhang, Yongquan AND Fan, Fenglei AND Elazab, Ahmed AND Wan, Xiang AND Wang, Changmiao},
        title = { { A Confounder-aware Representation Learning Framework for Alzheimer’s Disease Classification } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16886},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The main contribution is a causally inspired confounder-aware AD diagnosis framework that improves model robustness by explicitly reducing sensitivity to imaging artifacts and demographic biases through masking-based region analysis and covariate-conditioned feature modulation.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper is clearly written and well structured, making the proposed method easy to follow. The experimental evaluation is comprehensive and well designed, with appropriate comparisons that support the claims. In addition, the proposed modules are well motivated, with a clear connection to the problem of confounding factors in Alzheimer’s disease diagnosis, which strengthens the overall methodological design.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;A major weakness of the paper is that the ablation study suggests the main performance gain comes from the CAPM component rather than the MDM module, which is more prominently emphasized in the manuscript. In the comparison with state-of-the-art methods, it is also observed that two tabular-information-based models achieve strong performance, This creates a potential mismatch between the narrative emphasis and the empirical contributions. In particular, the specific advantage of CAPM over existing confounder-aware approaches is not clearly articulated, and the role of MDM appears less impactful than suggested in the main text. The paper would benefit from a clearer explanation of the distinct contribution of CAPM and a more balanced discussion of the actual empirical contributions of each module.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission has provided an anonymized link to the source code, dataset, or any other dependencies.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper is generally well written, clearly structured, and easy to follow. The proposed method is well motivated, and the experimental setup is comprehensive with meaningful comparisons. However, my concerns regarding the relative contributions of the proposed modules influenced my final evaluation. I rated the paper as a weak accept rather than a stronger accept.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.A prediction-sensitivity-based mask learning strategy for spatially robust representations. The paper proposes the Mask Decoupling Module (MDM), which estimates the importance of local image regions by masking individual patches and measuring the resulting change in model output. This provides a structured, perturbation-driven mechanism to separate predictive regions from spurious ones, encouraging the network to allocate representational capacity toward disease-relevant brain structures rather than incidental imaging patterns.
2.A covariate-conditioned feature modulation module for confounder suppression. The paper introduces the Confounder-Aware Patch Modulation (CAPM) module, which integrates tabular clinical and demographic variables to explicitly adjust patch-level visual features. By learning a per-patch scaling factor conditioned on covariates such as age and sex, the module provides a direct mechanism to attenuate variation in image representations that is attributable to confounding factors rather than disease status, promoting more robust multimodal fusion.
3.A unified, end-to-end multimodal framework inspired by causal reasoning for AD diagnosis. Rather than relying on multi-stage pipelines or post-hoc correction, the paper combines MDM and CAPM into a single jointly trained framework with a composite loss function that simultaneously encourages reliance on causal features, penalizes dependence on non-causal image regions, and prevents the tabular modality from dominating the prediction. This design offers a structurally simple yet principled approach to learning confounder-aware representations from heterogeneous data sources.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Principled problem formulation grounded in causal reasoning. The paper clearly articulates the confounding problem in AD diagnosis using a causal dependency diagram (Z → X, Z → Y), and designs each module to address a specific edge in the causal graph. This structured motivation, with MDM targeting spurious image regions and CAPM targeting covariate-linked variation, provides a coherent and interpretable framework, even if the implementation does not perform strict causal inference.
2.Joint handling of both visible and hidden confounders in a single end-to-end framework. Unlike prior methods that address either imaging artifacts or covariate confounding in isolation, this framework tackles both within a unified, end-to-end trainable pipeline. The adversarial losses on the non-causal branch and the tabular branch (maximizing their classification errors) provide an elegant mechanism to simultaneously suppress two distinct sources of spurious correlation without requiring multi-stage training.
3.Strong empirical performance with consistent gains across two independent datasets. The model achieves substantial improvements over the strongest baseline on ADNI (+10.4% ACC, +16.5% REC, +18.3% F1) and also reaches state-of-the-art on the external NACC dataset. Consistent improvement across both in-distribution and cross-cohort settings suggests that the framework does learn more generalizable representations, despite the architectural concerns raised above.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Architectural mismatch between CNN and patch-level causal reasoning. The MDM is designed to assess the contribution of individual patches to the prediction. However, the framework uses a CNN (ResNet) as the backbone, where deeper-layer features aggregate information across the entire spatial extent through hierarchical convolutions. This means that masking a single patch in the input does not cleanly isolate that patch’s contribution at the feature level, as the effect propagates non-linearly across spatially overlapping receptive fields. A Vision Transformer (ViT), which operates on explicit patch tokens with self-attention, would be a more architecturally consistent choice for patch-level importance estimation.&lt;/p&gt;

      &lt;p&gt;2.Questionable validity of the mean-image prior. The ranking-consistency loss uses patch scores from the training set’s voxel-wise mean image as a structural prior. However, averaging across all subjects simultaneously smooths out both disease-relevant anatomical variations (e.g., heterogeneous hippocampal atrophy patterns) and confounding noise. It is unclear whether the resulting patch ranking reflects diagnostically meaningful regions or merely morphologically stable areas across the population. Using an ambiguous supervisory signal to guide mask learning may misdirect the model.&lt;/p&gt;

      &lt;p&gt;3.Per-patch modulation in CAPM ignores global spatial dependencies. The modulation coefficient γ is computed independently for each patch, assuming that confounders affect patches in isolation. In reality, confounding factors such as aging-related brain atrophy have spatially correlated, global effects across the brain. Independent per-patch adjustment fails to capture these inter-region dependencies, potentially producing spatially incoherent adjusted features.
Scale of γ is unconstrained, and subsequent Pool3D undermines the modulation. The paper does not explicitly constrain the range of γ. If γ approaches or exceeds 1, the adjusted features (1−γ)⋅X(1 - \gamma) \cdot X
(1−γ)⋅X can collapse to near-zero or become negative, causing numerical instability. More critically, the adjusted features are immediately passed through a Pool3D layer, which aggregates across the spatial dimensions indiscriminately. This pooling blends heavily suppressed patches with minimally affected ones, effectively diluting the selective suppression that CAPM is designed to achieve. The modulation and the subsequent aggregation work against each other.&lt;/p&gt;

      &lt;p&gt;4.Lack of clarity on the multi-class AUC computation. The task is a three-class classification (NC vs. MCI vs. AD), yet the paper reports AUC without specifying the computation strategy (e.g., one-vs-rest macro or weighted averaging, or one-vs-one). Given the substantial class imbalance in both datasets (e.g., ADNI has NC:MCI:AD = 355:303:569), the choice of averaging scheme can significantly affect the reported numbers. This omission limits reproducibility and fair comparison.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission has provided an anonymized link to the source code, dataset, or any other dependencies.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper addresses a meaningful problem, learning confounder-aware representations for AD diagnosis, and the causal framing provides a coherent motivation. The dual-module design (MDM + CAPM) and the composite adversarial loss are conceptually appealing, and the empirical results show consistent improvements over baselines on two public datasets. However, several concerns regarding the core module designs prevent me from recommending acceptance at this stage.&lt;/p&gt;

      &lt;p&gt;1.The MDM relies on patch-level perturbation to estimate regional importance, yet the backbone is a CNN (ResNet) whose hierarchical convolutions cause deep features to aggregate information well beyond individual patch boundaries. Masking a single patch at the input level does not produce a spatially isolated effect at the feature level, which undermines the theoretical justification for treating patch-wise scores as meaningful estimates of local causal relevance. A patch-native architecture such as a Vision Transformer would be more consistent with this design.&lt;/p&gt;

      &lt;p&gt;2.The mean-image prior used to supervise mask learning is not well justified. Averaging across all training subjects smooths out both disease-related variation and confounding noise, making it unclear whether the resulting ranking signal reflects diagnostically informative regions or simply morphologically stable areas.&lt;/p&gt;

      &lt;p&gt;3.The CAPM module computes a per-patch modulation coefficient γ independently, ignoring the fact that confounders like age-related atrophy have spatially correlated global effects. Moreover, γ is not explicitly constrained in range, and the modulated features are immediately aggregated by Pool3D, which indiscriminately blends heavily suppressed and minimally affected patches. This effectively dilutes the selective suppression that CAPM is intended to achieve, creating a tension between the modulation step and the subsequent aggregation.&lt;/p&gt;

      &lt;p&gt;4.The paper reports AUC for a three-class task without specifying the computation strategy (one-vs-rest, one-vs-one, macro or weighted averaging), which limits reproducibility given the notable class imbalance in both datasets.&lt;/p&gt;

      &lt;p&gt;In summary, while the motivation is sound and the results are promising, the gap between the causal reasoning framework and the architectural choices that implement it raises fundamental questions about whether the modules function as claimed. I would be open to reconsidering if the authors can provide convincing explanations or additional experiments addressing these concerns in the rebuttal.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Reject&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper presents a genuinely interesting causal framing for AD diagnosis and demonstrates consistent empirical gains across two independent datasets, which are meaningful contributions that work in the authors’ favor; however, the reviewer’s core concern about the architectural mismatch between the CNN backbone and the patch-level causal reasoning was only partially addressed, and while the ViT ablation result was a welcome addition, it does not fully resolve the theoretical inconsistency embedded in the MDM design; the rebuttal on CAPM was also somewhat underdeveloped, particularly regarding the Pool3D aggregation issue and the unconstrained range of γ, where the authors’ explanation relied heavily on the claim of a misunderstanding without providing sufficient clarifying evidence; on the positive side, the strong cross-cohort generalization results on NACC and the largely satisfactory responses to the prior loss and AUC concerns do leave a realistic window for acceptance; on balance however, the unresolved architectural tensions in the two core modules remain the primary obstacle, and the rebuttal did not provide enough theoretical or empirical grounding to fully close the gap between the causal reasoning motivation and the actual implementation choices made in the paper.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents a multimodal, confounder-aware representation learning framework for AD classification. The MDM module identifies and prioritizes disease prediction relevant patches and CAPM weights patches based on tabular covariates, which are useful for achieving high performance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The motivation is well-formulated in the context of localized disease information, spurious correlations and confounder effects in brain images for disease diagnosis. The results are strong over baselines, the ablation is included, and the attention maps are biologically reasonable.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;First, the paper is a little over-emphasing the causal terminologies, while the method is more related to perturbation-based feature importance and prediction sensitivity, not causal effects. Second, many baselines are vision-only methods (CausalMixNet, CAAM, MAD-Former), so they may get disadvantages without some key tabular covariates like clinical assessments that highly correlate with diagnosis. Third, in the CAPM module, it might be simplistic to decrease weights for particular patches uniformly across all dimensions, since different features could have different correlations with confounders. Fourth, MDM module doesn’t seem to contribute much to the performance in the ablation (wo MDM’s F1 and AUC are close to full model). Fifth, the population voxel mean image as a prior may discourage the identification of individual-level disease-related variations.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission has provided an anonymized link to the source code, dataset, or any other dependencies.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper presents a causality-inspired, technically sound, sophisticated multimodal framework for AD classification. However, it’s unclear whether it’s the tabular features themselves that contribute more to high performance or the CAPM module. Ablation on MDM also does not fully demonstrate importance of the module.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Although the justifications for CAPM and MDM modules are still not convincing, the paper has strong empirical evidence and interpretability presentations.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;We thank all reviewers for their comprehensive and constructive feedback on our paper, particularly regarding “module contributions” (R1 and R3), “implementation of CAPM” (R2 and R3), and “prior score loss” (R2 and R3). We first address three general issues and then provide detailed responses to each reviewer.&lt;/p&gt;

  &lt;p&gt;G1: Module Contributions
MDM and CAPM address distinct issues: MDM penalizes unstable local dependencies via spatial regularization; CAPM adjusts for covariate-induced variation. When combined, MDM raises ACC from 0.819→0.833 (ADNI) and 0.802→0.831 (NACC); macro F1 and AUC also improve modestly (ADNI: F1 +0.012, AUC +0.009; NACC: F1 +0.010, AUC +0.008). Gains are modest but consistent, indicating CAPM is the primary performance driver, while MDM provides additional spatial regularization, interpretability. Because ADNI and NACC contain few severe artifacts or local noise, MDM’s marginal gain is expected. This complementary behavior aligns with our framework’s design rationale, where CAPM handles global covariate-related variation and MDM addresses local image instability.&lt;/p&gt;

  &lt;p&gt;G2: Implementation of CAPM 
We acknowledge confounders may have spatial dependencies. We clarify three points: (i) Each \gamma is computed from all confounders, enabling cross-patch learning, and acts on patch-level representations, not raw pixels. (ii) \gamma scale is unconstrained, but a residual connection prevents numerical instability. (iii) Pooling is applied per patch, preserving selective suppression; a global pooling claim is a misunderstanding. The current \gamma is a per-patch scalar (uniform across feature dimensions) for stable optimization;&lt;/p&gt;

  &lt;p&gt;G3: Prior Loss 
The mean image of the training set is not used for diagnostic supervision. Instead, the ranking-consistency loss acts as a weak, early-stage regularizer with small weight, linearly decaying to zero during training. It stabilizes initial mask learning, preventing degenerate or symmetric score distributions. In ablation experiments, removing this term led to larger variance in early training loss and slower convergence. The final masks are determined predominantly by the task-driven classification objective and perturbation-based MDM, ensuring that the ranking term does not bias final predictions.&lt;/p&gt;

  &lt;p&gt;R1Q1: Advantage
CAPM is not a simple tabular branch, it conditions local patch features on covariates. TabFormer alone achieves AUC 0.795 but ACC/F1 only 0.656/0.558 on ADNI, showing tabular data alone is insufficient. Replacing CAPM with FiLM/HyperFusion degrades performance, validating covariate-guided modulation. CAPM is the main empirical contributor; MDM provides complementary regularization and interpretability.&lt;/p&gt;

  &lt;p&gt;R2Q1: Architecture
MDM relies on perturbation-based sensitivity, not strict causal effect estimation. Despite CNN receptive field overlap, masking a patch selectively suppresses local activations; the resulting change in prediction still reflects the model’s reliance on that region, yielding a “blurred” but effective localization. This perturb-and-measure mechanism provides a useful regularization signal, as confirmed by our ablation results. We also tested a ViT backbone, which further improved performance, confirming the framework’s generality across architectures. We chose CNN for efficiency and fair baseline comparison; ViT ablation and discussion will be added to the manuscript.&lt;/p&gt;

  &lt;p&gt;R2Q4: Calculation of AUC 
Multi-class AUC uses one-vs-rest with macro averaging, giving equal weight to NC, MCI, AD to avoid majority-class bias.&lt;/p&gt;

  &lt;p&gt;R3Q1: Excessive Use of Causal Terminology 
We will replace “causal/non-causal branch” with “covariate-sensitive/non-sensitive branch” and rephrase causal claims as hypothesized dependencies.&lt;/p&gt;

  &lt;p&gt;R3Q2: Baseline Comparisons.
Image-only baselines (CausalMixNet, etc.) assess visual representation robustness. Tabular-only (TabFormer) and multimodal models (HyperFusion) ensure fair comparison. Our three-category baseline design gives a comprehensive evaluation.&lt;/p&gt;

&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Invite for Rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes a causality-inspired, confounder-aware multimodal framework (MDM + CAPM) for Alzheimer’s disease diagnosis. The reviews are mixed: R1 recommends weak accept, while R2 and R3 recommend weak reject. The ablation study shows that MDM’s independent contribution is limited and the main performance gain appears to come from CAPM, which creates a mismatch with the paper’s narrative emphasis and calls for a clearer, more balanced explanation; the validity of using a population-level voxel-wise mean image as a prior is also questioned as a meaningful supervisory signal. The authors are encouraged to focus their rebuttal on the attribution of each module’s contribution, as well as the specific concerns raised by the reviewers.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;During the rebuttal phase, the authors provided positive responses that resolved some detailed issues and addressed concerns regarding the prior score loss. R2 and R3 still have remaining concerns regarding the MDM CAPM. However, considering the comprehensive setup and interpretability, I lean towards acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-2&quot;&gt;Meta-review #2&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Reject&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The proposal is technically interesting but there is a large field of unbiased or confounder-aware learning that the authors did not consider in the comparison. The validation is mainly driven by the overall prediction accuracy, a metric that is irrelevant to the studied topic of confounder modeling. There is no evidence showing whether there are confounding effects in the prediction and whether those effects are captured and mitigated.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-3&quot;&gt;Meta-review #3&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The rebuttal reasonably clarified the relative roles of CAPM and MDM, the weak regularizing nature of the prior loss, the AUC computation, and the intended softening of causal claims; although some architectural concerns remain, the consistent empirical gains and improved positioning support acceptance with these clarifications incorporated in the camera-ready version.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Chen, Yuanhao AND Chen, Jingwen AND Xu, Jinyi AND Yu, Haoqi AND Wang, Bin AND Li, Chunzhong AND Zhang, Yongquan AND Fan, Fenglei AND Elazab, Ahmed AND Wan, Xiang AND Wang, Changmiao</name></author><category term="Body -&gt; Brain" /><category term="Modalities -&gt; MRI" /><category term="Applications -&gt; Computer-Aided Diagnosis" /><category term="Machine Learning -&gt; Deep Learning" /><category term="Machine Learning -&gt; Multimodal Models / LLMs / VLMs" /><category term="Chen, Yuanhao" /><category term="Chen, Jingwen" /><category term="Xu, Jinyi" /><category term="Yu, Haoqi" /><category term="Wang, Bin" /><category term="Li, Chunzhong" /><category term="Zhang, Yongquan" /><category term="Fan, Fenglei" /><category term="Elazab, Ahmed" /><category term="Wan, Xiang" /><category term="Wang, Changmiao" /><summary type="html">Abstract Alzheimer’s disease (AD) is the most prevalent neurodegenerative disorder, and current treatments cannot reverse its progression, making early diagnosis and intervention essential. Deep learning methods have shown promise for diagnosing AD from neuroimaging data, but these models are often affected by both visible and hidden sources of noise. Visible noise typically reflects systematic imaging differences, such as scanner-related biases and acquisition artifacts, whereas hidden noise is linked to individual variation that can lead models to depend on spurious patterns that do not generalize well. Most existing approaches emphasize statistical associations and do not adequately account for these confounding influences. Motivated by these limitations, we propose a multimodal, confounder-aware framework for AD diagnosis inspired by causal reasoning. The model includes a region-masking component (Mask Decoupling Module) that probes the image by making small, targeted changes to localized areas, then highlights the regions that most strongly affect the prediction, encouraging the network to prioritize structurally informative features. We also introduce a covariate-guided feature modulation mechanism that uses clinical and demographic variables to adjust local image features, thereby reducing variation attributable to these factors. Although the framework does not estimate causal effects directly, it is designed to learn more stable representations by limiting reliance on shortcuts, improving robustness and generalization in AD diagnostic models. Our code is available at \url{https://anonymous.4open.science/r/Causal_fusion-4E6E}. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/1600_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository https://github.com/redtea-code/Causal_fusion Link to the Dataset(s) N/A BibTex @InProceedings{CheYua_AConfounderaware_MICCAI2026,         author = { Chen, Yuanhao AND Chen, Jingwen AND Xu, Jinyi AND Yu, Haoqi AND Wang, Bin AND Li, Chunzhong AND Zhang, Yongquan AND Fan, Fenglei AND Elazab, Ahmed AND Wan, Xiang AND Wang, Changmiao},         title = { { A Confounder-aware Representation Learning Framework for Alzheimer’s Disease Classification } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16886},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper The main contribution is a causally inspired confounder-aware AD diagnosis framework that improves model robustness by explicitly reducing sensitivity to imaging artifacts and demographic biases through masking-based region analysis and covariate-conditioned feature modulation. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The paper is clearly written and well structured, making the proposed method easy to follow. The experimental evaluation is comprehensive and well designed, with appropriate comparisons that support the claims. In addition, the proposed modules are well motivated, with a clear connection to the problem of confounding factors in Alzheimer’s disease diagnosis, which strengthens the overall methodological design. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. A major weakness of the paper is that the ablation study suggests the main performance gain comes from the CAPM component rather than the MDM module, which is more prominently emphasized in the manuscript. In the comparison with state-of-the-art methods, it is also observed that two tabular-information-based models achieve strong performance, This creates a potential mismatch between the narrative emphasis and the empirical contributions. In particular, the specific advantage of CAPM over existing confounder-aware approaches is not clearly articulated, and the role of MDM appears less impactful than suggested in the main text. The paper would benefit from a clearer explanation of the distinct contribution of CAPM and a more balanced discussion of the actual empirical contributions of each module. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission has provided an anonymized link to the source code, dataset, or any other dependencies. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The paper is generally well written, clearly structured, and easy to follow. The proposed method is well motivated, and the experimental setup is comprehensive with meaningful comparisons. However, my concerns regarding the relative contributions of the proposed modules influenced my final evaluation. I rated the paper as a weak accept rather than a stronger accept. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #2 Please describe the contribution of the paper 1.A prediction-sensitivity-based mask learning strategy for spatially robust representations. The paper proposes the Mask Decoupling Module (MDM), which estimates the importance of local image regions by masking individual patches and measuring the resulting change in model output. This provides a structured, perturbation-driven mechanism to separate predictive regions from spurious ones, encouraging the network to allocate representational capacity toward disease-relevant brain structures rather than incidental imaging patterns. 2.A covariate-conditioned feature modulation module for confounder suppression. The paper introduces the Confounder-Aware Patch Modulation (CAPM) module, which integrates tabular clinical and demographic variables to explicitly adjust patch-level visual features. By learning a per-patch scaling factor conditioned on covariates such as age and sex, the module provides a direct mechanism to attenuate variation in image representations that is attributable to confounding factors rather than disease status, promoting more robust multimodal fusion. 3.A unified, end-to-end multimodal framework inspired by causal reasoning for AD diagnosis. Rather than relying on multi-stage pipelines or post-hoc correction, the paper combines MDM and CAPM into a single jointly trained framework with a composite loss function that simultaneously encourages reliance on causal features, penalizes dependence on non-causal image regions, and prevents the tabular modality from dominating the prediction. This design offers a structurally simple yet principled approach to learning confounder-aware representations from heterogeneous data sources. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.Principled problem formulation grounded in causal reasoning. The paper clearly articulates the confounding problem in AD diagnosis using a causal dependency diagram (Z → X, Z → Y), and designs each module to address a specific edge in the causal graph. This structured motivation, with MDM targeting spurious image regions and CAPM targeting covariate-linked variation, provides a coherent and interpretable framework, even if the implementation does not perform strict causal inference. 2.Joint handling of both visible and hidden confounders in a single end-to-end framework. Unlike prior methods that address either imaging artifacts or covariate confounding in isolation, this framework tackles both within a unified, end-to-end trainable pipeline. The adversarial losses on the non-causal branch and the tabular branch (maximizing their classification errors) provide an elegant mechanism to simultaneously suppress two distinct sources of spurious correlation without requiring multi-stage training. 3.Strong empirical performance with consistent gains across two independent datasets. The model achieves substantial improvements over the strongest baseline on ADNI (+10.4% ACC, +16.5% REC, +18.3% F1) and also reaches state-of-the-art on the external NACC dataset. Consistent improvement across both in-distribution and cross-cohort settings suggests that the framework does learn more generalizable representations, despite the architectural concerns raised above. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.Architectural mismatch between CNN and patch-level causal reasoning. The MDM is designed to assess the contribution of individual patches to the prediction. However, the framework uses a CNN (ResNet) as the backbone, where deeper-layer features aggregate information across the entire spatial extent through hierarchical convolutions. This means that masking a single patch in the input does not cleanly isolate that patch’s contribution at the feature level, as the effect propagates non-linearly across spatially overlapping receptive fields. A Vision Transformer (ViT), which operates on explicit patch tokens with self-attention, would be a more architecturally consistent choice for patch-level importance estimation. 2.Questionable validity of the mean-image prior. The ranking-consistency loss uses patch scores from the training set’s voxel-wise mean image as a structural prior. However, averaging across all subjects simultaneously smooths out both disease-relevant anatomical variations (e.g., heterogeneous hippocampal atrophy patterns) and confounding noise. It is unclear whether the resulting patch ranking reflects diagnostically meaningful regions or merely morphologically stable areas across the population. Using an ambiguous supervisory signal to guide mask learning may misdirect the model. 3.Per-patch modulation in CAPM ignores global spatial dependencies. The modulation coefficient γ is computed independently for each patch, assuming that confounders affect patches in isolation. In reality, confounding factors such as aging-related brain atrophy have spatially correlated, global effects across the brain. Independent per-patch adjustment fails to capture these inter-region dependencies, potentially producing spatially incoherent adjusted features. Scale of γ is unconstrained, and subsequent Pool3D undermines the modulation. The paper does not explicitly constrain the range of γ. If γ approaches or exceeds 1, the adjusted features (1−γ)⋅X(1 - \gamma) \cdot X (1−γ)⋅X can collapse to near-zero or become negative, causing numerical instability. More critically, the adjusted features are immediately passed through a Pool3D layer, which aggregates across the spatial dimensions indiscriminately. This pooling blends heavily suppressed patches with minimally affected ones, effectively diluting the selective suppression that CAPM is designed to achieve. The modulation and the subsequent aggregation work against each other. 4.Lack of clarity on the multi-class AUC computation. The task is a three-class classification (NC vs. MCI vs. AD), yet the paper reports AUC without specifying the computation strategy (e.g., one-vs-rest macro or weighted averaging, or one-vs-one). Given the substantial class imbalance in both datasets (e.g., ADNI has NC:MCI:AD = 355:303:569), the choice of averaging scheme can significantly affect the reported numbers. This omission limits reproducibility and fair comparison. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission has provided an anonymized link to the source code, dataset, or any other dependencies. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The paper addresses a meaningful problem, learning confounder-aware representations for AD diagnosis, and the causal framing provides a coherent motivation. The dual-module design (MDM + CAPM) and the composite adversarial loss are conceptually appealing, and the empirical results show consistent improvements over baselines on two public datasets. However, several concerns regarding the core module designs prevent me from recommending acceptance at this stage. 1.The MDM relies on patch-level perturbation to estimate regional importance, yet the backbone is a CNN (ResNet) whose hierarchical convolutions cause deep features to aggregate information well beyond individual patch boundaries. Masking a single patch at the input level does not produce a spatially isolated effect at the feature level, which undermines the theoretical justification for treating patch-wise scores as meaningful estimates of local causal relevance. A patch-native architecture such as a Vision Transformer would be more consistent with this design. 2.The mean-image prior used to supervise mask learning is not well justified. Averaging across all training subjects smooths out both disease-related variation and confounding noise, making it unclear whether the resulting ranking signal reflects diagnostically informative regions or simply morphologically stable areas. 3.The CAPM module computes a per-patch modulation coefficient γ independently, ignoring the fact that confounders like age-related atrophy have spatially correlated global effects. Moreover, γ is not explicitly constrained in range, and the modulated features are immediately aggregated by Pool3D, which indiscriminately blends heavily suppressed and minimally affected patches. This effectively dilutes the selective suppression that CAPM is intended to achieve, creating a tension between the modulation step and the subsequent aggregation. 4.The paper reports AUC for a three-class task without specifying the computation strategy (one-vs-rest, one-vs-one, macro or weighted averaging), which limits reproducibility given the notable class imbalance in both datasets. In summary, while the motivation is sound and the results are promising, the gap between the causal reasoning framework and the architectural choices that implement it raises fundamental questions about whether the modules function as claimed. I would be open to reconsidering if the authors can provide convincing explanations or additional experiments addressing these concerns in the rebuttal. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Reject [Post rebuttal] Please justify your final decision from above. The paper presents a genuinely interesting causal framing for AD diagnosis and demonstrates consistent empirical gains across two independent datasets, which are meaningful contributions that work in the authors’ favor; however, the reviewer’s core concern about the architectural mismatch between the CNN backbone and the patch-level causal reasoning was only partially addressed, and while the ViT ablation result was a welcome addition, it does not fully resolve the theoretical inconsistency embedded in the MDM design; the rebuttal on CAPM was also somewhat underdeveloped, particularly regarding the Pool3D aggregation issue and the unconstrained range of γ, where the authors’ explanation relied heavily on the claim of a misunderstanding without providing sufficient clarifying evidence; on the positive side, the strong cross-cohort generalization results on NACC and the largely satisfactory responses to the prior loss and AUC concerns do leave a realistic window for acceptance; on balance however, the unresolved architectural tensions in the two core modules remain the primary obstacle, and the rebuttal did not provide enough theoretical or empirical grounding to fully close the gap between the causal reasoning motivation and the actual implementation choices made in the paper. Review #3 Please describe the contribution of the paper This paper presents a multimodal, confounder-aware representation learning framework for AD classification. The MDM module identifies and prioritizes disease prediction relevant patches and CAPM weights patches based on tabular covariates, which are useful for achieving high performance. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The motivation is well-formulated in the context of localized disease information, spurious correlations and confounder effects in brain images for disease diagnosis. The results are strong over baselines, the ablation is included, and the attention maps are biologically reasonable. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. First, the paper is a little over-emphasing the causal terminologies, while the method is more related to perturbation-based feature importance and prediction sensitivity, not causal effects. Second, many baselines are vision-only methods (CausalMixNet, CAAM, MAD-Former), so they may get disadvantages without some key tabular covariates like clinical assessments that highly correlate with diagnosis. Third, in the CAPM module, it might be simplistic to decrease weights for particular patches uniformly across all dimensions, since different features could have different correlations with confounders. Fourth, MDM module doesn’t seem to contribute much to the performance in the ablation (wo MDM’s F1 and AUC are close to full model). Fifth, the population voxel mean image as a prior may discourage the identification of individual-level disease-related variations. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission has provided an anonymized link to the source code, dataset, or any other dependencies. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The paper presents a causality-inspired, technically sound, sophisticated multimodal framework for AD classification. However, it’s unclear whether it’s the tabular features themselves that contribute more to high performance or the CAPM module. Ablation on MDM also does not fully demonstrate importance of the module. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. Although the justifications for CAPM and MDM modules are still not convincing, the paper has strong empirical evidence and interpretability presentations. Author Feedback We thank all reviewers for their comprehensive and constructive feedback on our paper, particularly regarding “module contributions” (R1 and R3), “implementation of CAPM” (R2 and R3), and “prior score loss” (R2 and R3). We first address three general issues and then provide detailed responses to each reviewer. G1: Module Contributions MDM and CAPM address distinct issues: MDM penalizes unstable local dependencies via spatial regularization; CAPM adjusts for covariate-induced variation. When combined, MDM raises ACC from 0.819→0.833 (ADNI) and 0.802→0.831 (NACC); macro F1 and AUC also improve modestly (ADNI: F1 +0.012, AUC +0.009; NACC: F1 +0.010, AUC +0.008). Gains are modest but consistent, indicating CAPM is the primary performance driver, while MDM provides additional spatial regularization, interpretability. Because ADNI and NACC contain few severe artifacts or local noise, MDM’s marginal gain is expected. This complementary behavior aligns with our framework’s design rationale, where CAPM handles global covariate-related variation and MDM addresses local image instability. G2: Implementation of CAPM We acknowledge confounders may have spatial dependencies. We clarify three points: (i) Each \gamma is computed from all confounders, enabling cross-patch learning, and acts on patch-level representations, not raw pixels. (ii) \gamma scale is unconstrained, but a residual connection prevents numerical instability. (iii) Pooling is applied per patch, preserving selective suppression; a global pooling claim is a misunderstanding. The current \gamma is a per-patch scalar (uniform across feature dimensions) for stable optimization; G3: Prior Loss The mean image of the training set is not used for diagnostic supervision. Instead, the ranking-consistency loss acts as a weak, early-stage regularizer with small weight, linearly decaying to zero during training. It stabilizes initial mask learning, preventing degenerate or symmetric score distributions. In ablation experiments, removing this term led to larger variance in early training loss and slower convergence. The final masks are determined predominantly by the task-driven classification objective and perturbation-based MDM, ensuring that the ranking term does not bias final predictions. R1Q1: Advantage CAPM is not a simple tabular branch, it conditions local patch features on covariates. TabFormer alone achieves AUC 0.795 but ACC/F1 only 0.656/0.558 on ADNI, showing tabular data alone is insufficient. Replacing CAPM with FiLM/HyperFusion degrades performance, validating covariate-guided modulation. CAPM is the main empirical contributor; MDM provides complementary regularization and interpretability. R2Q1: Architecture MDM relies on perturbation-based sensitivity, not strict causal effect estimation. Despite CNN receptive field overlap, masking a patch selectively suppresses local activations; the resulting change in prediction still reflects the model’s reliance on that region, yielding a “blurred” but effective localization. This perturb-and-measure mechanism provides a useful regularization signal, as confirmed by our ablation results. We also tested a ViT backbone, which further improved performance, confirming the framework’s generality across architectures. We chose CNN for efficiency and fair baseline comparison; ViT ablation and discussion will be added to the manuscript. R2Q4: Calculation of AUC Multi-class AUC uses one-vs-rest with macro averaging, giving equal weight to NC, MCI, AD to avoid majority-class bias. R3Q1: Excessive Use of Causal Terminology We will replace “causal/non-causal branch” with “covariate-sensitive/non-sensitive branch” and rephrase causal claims as hypothesized dependencies. R3Q2: Baseline Comparisons. Image-only baselines (CausalMixNet, etc.) assess visual representation robustness. Tabular-only (TabFormer) and multimodal models (HyperFusion) ensure fair comparison. Our three-category baseline design gives a comprehensive evaluation. Meta-Review Meta-review #1 Your recommendation Invite for Rebuttal Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. This paper proposes a causality-inspired, confounder-aware multimodal framework (MDM + CAPM) for Alzheimer’s disease diagnosis. The reviews are mixed: R1 recommends weak accept, while R2 and R3 recommend weak reject. The ablation study shows that MDM’s independent contribution is limited and the main performance gain appears to come from CAPM, which creates a mismatch with the paper’s narrative emphasis and calls for a clearer, more balanced explanation; the validity of using a population-level voxel-wise mean image as a prior is also questioned as a meaningful supervisory signal. The authors are encouraged to focus their rebuttal on the attribution of each module’s contribution, as well as the specific concerns raised by the reviewers. After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. During the rebuttal phase, the authors provided positive responses that resolved some detailed issues and addressed concerns regarding the prior score loss. R2 and R3 still have remaining concerns regarding the MDM CAPM. However, considering the comprehensive setup and interpretability, I lean towards acceptance. Meta-review #2 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Reject Please justify your recommendation. The proposal is technically interesting but there is a large field of unbiased or confounder-aware learning that the authors did not consider in the comparison. The validation is mainly driven by the overall prediction accuracy, a metric that is irrelevant to the studied topic of confounder modeling. There is no evidence showing whether there are confounding effects in the prediction and whether those effects are captured and mitigated. Meta-review #3 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. The rebuttal reasonably clarified the relative roles of CAPM and MDM, the weak regularizing nature of the prior loss, the AUC computation, and the intended softening of causal claims; although some architectural concerns remain, the consistent empirical gains and improved positioning support acceptance with these clarifications incorporated in the camera-ready version. back to top</summary></entry><entry><title type="html">A Counterfactual Framework for Directional Cell–Cell Interaction Analysis in Spatial Transcriptomics</title><link href="https://papers.miccai.org/miccai-2026/0006-Paper6105" rel="alternate" type="text/html" title="A Counterfactual Framework for Directional Cell–Cell Interaction Analysis in Spatial Transcriptomics" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0006-Paper6105</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0006-Paper6105">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;Understanding neighboring cell influences is central to spatial transcriptomics, yet existing methods rely on correlation or predefined ligand–receptor (LR) pairs and do not test directionality. We introduce a counterfactual, intervention-based framework for inferring directional cell–cell influence that is LR-agnostic and tests sender specificity. A neighborhood graph model predicts receiver cell state from local spatial context. Directional influence is quantified by counterfactually replacing neighbors of a sender type and measuring resulting displacement in predicted receiver state. We define a Counterfactual Directionality Score (CDS) that quantifies directional influence, and compute pairwise CDS by aggregating across receiver cells and test cores sender–receiver pair. Applied to Xenium cholangiocarcinoma tissue microarrays (38 cores), the framework identified reproducible, asymmetric interactions between tumor, immune, and stromal compartments; e.g. Tumor-EMT →Macrophage (CDS: 0.0828) and Fibroblast→Macrophage (CDS: 0.0582). Effects exceeded label-permutation and spatial-shuffle null models (p &amp;lt; 0.001, FDR-controlled) and remained stable under core-level bootstrap resampling. Inferred directional strengths correlated strongly with matched LR scores (r = 0.758, p = 0.0027), supporting biological concordance. Code:https://github.com/Snigdha022/Counterfactual-Cell-Influence-Spatial-Transcriptomics.git.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/6105_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/6105_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/Snigdha022/Counterfactual-Cell-Influence-Spatial-Transcriptomics.git&quot;&gt;https://github.com/Snigdha022/Counterfactual-Cell-Influence-Spatial-Transcriptomics.git&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;We utilized spatial transcriptomics data from human cholangiocarcinoma TMAs, selecting 38 independent cores generated via the Xenium platform. These data were obtained from The University of Texas MD Anderson Cancer Center and are not publicly available at this time. The code is demonstrated using a publicly available breast cancer dataset with the following link: &lt;a href=&quot;https://github.com/Snigdha022/Counterfactual-Cell-Influence-Spatial-Transcriptomics/tree/main/data/raw&quot;&gt;https://github.com/Snigdha022/Counterfactual-Cell-Influence-Spatial-Transcriptomics/tree/main/data/raw&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{AnzHum_ACounterfactual_MICCAI2026,
        author = { Anzum, Humaira AND Kochat, Veena AND Mahmud, Md Ishtyaq AND Dwarampudi, Jagan Mohan Reddy AND Satpati, Suresh AND Shukla, Pooja AND Javle, Milind AND Kwong, Lawrence AND Rai, Kunal AND Banerjee, Tania},
        title = { { A Counterfactual Framework for Directional Cell–Cell Interaction Analysis in Spatial Transcriptomics } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16891},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes a novel counterfactual, intervention-based framework for inferring directional cell–cell interactions in spatial transcriptomics. Unlike conventional approaches that rely on correlation or predefined ligand–receptor (LR) priors, the method introduces the Counterfactual Directionality Score (CDS), which quantifies directional influence by explicitly perturbing the sender cell population and measuring the resulting change in the predicted receiver cell state. This formulation enables direct testing of causal directionality in a data-driven manner. The framework is further supported by a neighborhood-conditioned predictive model, rigorous statistical validation using multiple null models with FDR control, and biological validation via concordance with known LR interactions. Overall, the work presents a principled and scalable approach for directional cell–cell communication analysis.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Methodological novelty through counterfactual inference&lt;/p&gt;

      &lt;p&gt;The proposed framework represents a clear departure from conventional correlation-based or LR-dependent approaches. By introducing explicit counterfactual interventions (e. g. , sender type swapping) and defining CDS, the method directly tests directional influence rather than inferring it indirectly. This is a conceptually significant advancement in spatial transcriptomics analysis.&lt;/p&gt;

      &lt;p&gt;2.Strong statistical rigor and experimental design&lt;/p&gt;

      &lt;p&gt;The study employs a carefully structured evaluation pipeline, including strict separation of TMA cores into training, validation, and test sets. The use of multiple complementary null models (label permutation, spatial shuffling, and sender-agnostic replacement) with FDR correction ensures robustness against spurious findings. Additionally, core-level bootstrap resampling provides reliable uncertainty quantification.&lt;/p&gt;

      &lt;p&gt;3.Biological relevance and interpretability&lt;/p&gt;

      &lt;p&gt;Despite being fully data-driven, the inferred interactions show strong concordance with established biological knowledge, as evidenced by a high correlation with LR-based scores. The projection of gene-level changes onto functional programs further enhances interpretability and provides meaningful biological insights that are valuable for both basic and translational research.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Limited granularity in program-level interpretation&lt;/p&gt;

      &lt;p&gt;While the program-level projection provides a useful summary, the analysis lacks detailed identification of the genes driving each functional program shift. As a result, the biological interpretation remains somewhat high-level and less mechanistically grounded.&lt;/p&gt;

      &lt;p&gt;Suggestion: Reporting top contributing genes and their associated pathways for each program-level effect would significantly improve interpretability.&lt;/p&gt;

      &lt;p&gt;2.Insufficient analysis of neighborhood modeling choices&lt;/p&gt;

      &lt;p&gt;The use of a fixed number of nearest neighbors (k=20) and a specific exponential distance weighting function is not sufficiently justified or analyzed. The sensitivity of the results to these design choices remains unclear.&lt;/p&gt;

      &lt;p&gt;Suggestion: Performing sensitivity analyses over different values of k and alternative weighting schemes would strengthen the robustness claims. Incorporating more biologically grounded approaches (e. g. , optimal transport-based methods such as COMMOT) could further enhance realism.&lt;/p&gt;

      &lt;p&gt;3.Lack of clarity in counterfactual replacement procedure&lt;/p&gt;

      &lt;p&gt;The description of how spatial structure is preserved during counterfactual replacement (e. g. , “distance bin structure”) is not sufficiently detailed. This limits reproducibility and raises potential concerns regarding implementation.&lt;/p&gt;

      &lt;p&gt;Suggestion: Providing pseudocode or a more explicit algorithmic description would improve transparency and reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This is a well-motivated and carefully executed study that introduces a meaningful conceptual advance in spatial transcriptomics. The counterfactual formulation is both elegant and practically useful, and the experimental validation is rigorous. To further strengthen the work, the authors are encouraged to improve the granularity of biological interpretation and provide additional clarity on key methodological components, particularly the counterfactual replacement procedure and neighborhood modeling assumptions. Addressing these points would enhance both interpretability and reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The primary factors supporting this recommendation are the originality of the proposed counterfactual framework and the strong statistical rigor of the experimental design. The work addresses an important limitation in current spatial transcriptomics methodologies and provides a principled solution that enables explicit testing of directional influence. The results are supported by robust validation strategies and demonstrate meaningful biological consistency. While there are some limitations in interpretability detail and modeling choices, these are relatively minor compared to the overall contribution. Therefore, the paper meets the standard for acceptance at MICCAI.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The rebuttal satisfactorily addresses the main methodological and interpretational concerns raised during review. In particular, the authors appropriately agree to reframe the method as interventional model interpretability rather than formal causal inference, and to revise the causal language accordingly. This revised framing is important because it better aligns the contribution with the actual experiments: a controlled counterfactual perturbation framework for analyzing directional effects in the model-predicted receiver-cell state. The authors also clarify the counterfactual replacement procedure, specifying same-core replacement while preserving distance-bin structure, with fallback rules when exact matches are unavailable. This substantially reduces ambiguity in the implementation and improves reproducibility. The additional clarifications on biological interpretation and neighborhood modeling assumptions are also sufficient for the scope of this work. Some limitations remain, including evaluation on a single cancer type, a limited number of held-out test cores, and the lack of direct numerical comparison with methods (e.g., CellChat, NicheNet, or GITIII). However, I do not consider these limitations fatal. The paper already provides meaningful validation through matched null models, FDR control, bootstrap uncertainty estimation, and ligand–receptor concordance, and the authors reasonably explain that CDS produces a displacement-based counterfactual score that is not directly comparable to ligand–receptor co-expression scores or attention-based attribution outputs.&lt;/p&gt;

      &lt;p&gt;But reproducibility remains an important issue for the impact of this work. My accept-side recommendation assumes that the authors will fulfill their rebuttal commitment to publicly release the source code and relevant data or dependencies upon acceptance. Given the methodological nature of the contribution and the remaining implementation details, code release is important for verifying reproducibility, enabling independent use of the framework, and increasing the paper’s practical impact. I also strongly recommend that the authors carefully revise the manuscript presentation in the camera-ready version. In addition to removing causal overclaims, the paper would benefit from a more formal and polished structure, clearer methodological exposition, and correction of formatting issues (e.g., including standardizing the abstract format as a single coherent paragraph and fixing minor reference/figure-text inconsistencies or all the others (including reviewers’ comments) 100% completely).&lt;/p&gt;

      &lt;p&gt;Overall, the rebuttal resolves the main concerns to my satisfaction. The remaining issues are limitations and camera-ready responsibilities rather than grounds for rejection. I therefore keep my score with weak acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors present a counterfactual framework for inferring directional cell–cell influence in spatial transcriptomics without relying on predefined ligand–receptor priors. The experimental evaluation reveals reproducible and asymmetric directional influences between tumor, immune, and stromal compartments, most prominently Tumor-EMT→Macrophage and Fibroblast→Macrophage. Three complementary null models confirm that the observed effects are unlikely to arise by chance, and a ligand–receptor concordance analysis supports alignment with established biology (r = +0.758, p = 0.0027).&lt;/p&gt;

      &lt;p&gt;The main contribution of this work is the introduction of a statistically rigorous, intervention-based alternative to correlation-driven cell–cell communication analysis, applicable to large imaging transcriptomics datasets and clinical samples. Central to this contribution is the Counterfactual Directionality Score (CDS), a metric that quantifies directional influence by counterfactually replacing sender cells within spatial neighborhoods and measuring the resulting displacement in the predicted receiver cell state. Together, these elements constitute a principled and scalable framework for dissecting directional signaling in the tumor microenvironment.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;The strong concordance between the inferred counterfactual directionality scores and established ligand–receptor scores (r = +0.758, p = 0.0027), which provides meaningful external biological validation and supports the interpretation that the framework captures genuine signaling relationships rather than statistical artifacts.&lt;/li&gt;
        &lt;li&gt;The adoption of an explicit interventional framework that, unlike conventional correlation-based approaches, enables directional causal testing of cell–cell communication in spatial transcriptomics.&lt;/li&gt;
        &lt;li&gt;Ligand–receptor agnostic design, which eliminates dependence on curated interaction databases that are inherently incomplete and subject to annotation bias, thereby broadening the applicability of the framework to less-characterized biological contexts.&lt;/li&gt;
        &lt;li&gt;The rigorous statistical validation strategy, which combines three complementary null models with FDR correction to confirm that observed effects are unlikely to arise by chance. This is further supported by core-level bootstrap resampling and strict train/test separation, together providing robust evidence for the stability and reproducibility of the reported findings.&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;A limitation of the biological interpretation layer is that the projection of counterfactual effects onto functional programs is inherently contingent on the quality and completeness of the predefined gene sets used.&lt;/li&gt;
        &lt;li&gt;The experimental evaluation is the relatively small number of held-out test cores used for counterfactual analysis. With only 10 test cores drawn from a tissue microarray, the evaluation may not adequately capture the extent of inter-patient biological variability inherent to cholangiocarcinoma, a disease known for its considerable tumor microenvironment heterogeneity.&lt;/li&gt;
        &lt;li&gt;A minor but clear limitation of the paper is that the proposed framework is evaluated exclusively on a single cancer type (cholangiocarcinoma) which restricts the extent to which the findings and the method’s performance can be generalized to other tissue types, tumor microenvironments, or disease contexts. This could be part of the on going or future work.&lt;/li&gt;
      &lt;/ul&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The text reports the MSE values in panel A of Figure 1 for true-NI (0.849), MLP (1.032), and shuffled-NI (0.964), however, the corresponding values visible in the figure are 0.846, 1.036, and 0.969 respectively, indicating a discrepancy between the reported numbers and the actual figure.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(5) Accept — should be accepted, independent of rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;A statistically rigorous, intervention-based alternative to correlation-driven cell–cell communication analysis, applicable to large imaging transcriptomics datasets and clinical samples. Central to this contribution is the Counterfactual Directionality Score (CDS), a metric that quantifies directional influence by counterfactually replacing sender cells within spatial neighborhoods and measuring the resulting displacement in the predicted receiver cell state. Together, these elements constitute a principled and scalable framework for dissecting directional signaling in the tumor microenvironment. In addition, the adoption of an explicit interventional framework that, unlike conventional correlation-based approaches, enables directional causal testing of cell–cell communication in spatial transcriptomics. Finally, LR agnostic design, which eliminates dependence on curated interaction databases that are inherently incomplete and subject to annotation bias, thereby broadening the applicability of the framework to less-characterized biological contexts.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Somewhat confident (2)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Summary: The paper proposes a counterfactual framework to infer directional cell-cell influences in spatial transcriptomics data without the use of domain priors. The counterfactual framework involves (i) prediction of receiver cell state by a neighbourhood-conditioned graph model, (ii) measuring the displacement in predicted receiver state after controlled perturbations of sender composition.
The main claims of the paper are:
Claim 1: Existing approaches to inferring cell-cell influences are correlational and either depend on prior knowledge or do not explicitly test directionality. To bridge this gap, they propose a counterfactual framework.
Claim 2: The proposed counterfactual framework enables directional inference through intervention testing rather than association. 
Claim 3: Counterfactual Directionality score (CDS) disentangles cell type identity from intra-type heterogeneity.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The paper proposes an interesting approach for model interpretability rather than causal analysis. The proposed “CDS” score also seems to provide fine-grained interpretability information regarding cell identity and intra-type heterogeneity (Claim 3), as validated in their robustness analysis. It is also good to see the biological interpretation of the model artefacts.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The paper is jargon-heavy and hard to follow. Revision of some text and providing diagrams (For example: a block diagram for the methodology) would improve the readability of the paper. The motivation for the specific design choices of the methodology was missing. To name a few examples: Why was neighbourhood-conditioned graph model chosen? Has it been applied to similar problem. Why do we need to isolate functional state from cell type identity? Why were the two swap operators selected? It is also necessary to provide evidence from literature for these design choices.&lt;/p&gt;

      &lt;p&gt;2.Here, I am assessing claim 2 from the methodological perspective, i.e., the question - does the proposed framework enable directional inference of cell-cell communication within the data. To be able to infer the true directional cell-cell influences, the graph model adopted should be close to 100% accurate. Also, one of the main assumptions here is that there is no hidden confounding (measurement or environment induced artefacts). So, what is actually being achieved is model interpretability. Does node A influence node B for my model? If so, what is the direction of influence. The experiments are all going in this direction.&lt;/p&gt;

      &lt;p&gt;3.Following point 1, claim 1 is invalidated. The proposed methodology is still giving us predictive attribution and association.&lt;/p&gt;

      &lt;p&gt;4.Experiments comparing the proposed method with the existing methods mentioned in the introduction are missing. As, the proposed methodology is also giving some form of predictive attribution, it is important to assess how the framework performs in comparison to existing methods like CellChat, NicheNet, GITIII,…&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Poor&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Detailed constructive comments/action points:
1.The main premise of the paper and the claims need to be reformulated. It will be interesting to sell this as a model interpretability method rather than providing causal guarantees. 
2.Comparison to existing methods are necessary.
3.Rephrase the paper extensively for better readability and provide substantial evidence from literature. Also, clearly state all assumptions for the methodology.&lt;/p&gt;

      &lt;p&gt;Minor comment:
The doi link seem to appear twice for every reference in the paper.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(2) Reject — should be rejected, independent of rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The fundamental premise of the paper (claiming causal directional inference of cell-cell communications) is methodologically flawed. Even if the authors tweak the claims to frame this as a model interpretability tool, the manuscript still requires substantial revisions to clarify the methodology, as well as the addition of necessary SOTA baseline comparisons (e.g., CellChat, NicheNet) to prove its empirical value.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors agreed to modify the claims and state the assumptions explicitly. I am not convinced by the authors’ justification for lack of baseline comparisons, because direct numerical comparison is not required to answer the question “What novel cell-cell influences does the current method discover that the existing approaches miss?”. Despite this, I am recommending acceptance because of the biological perspectives offered in the paper, which is often missing from these kind of interdisciplinary computational papers.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;We thank all reviewers for their constructive feedback.
R#1:
Gene‑level interpretability: We agree that reporting the specific genes driving each program improves interpretability. We identified the top three up- and down-regulated genes per program for each significant sender–receiver pair using the same model and data. For example, in Fibroblast→Macrophage, Macrophage_Activation is driven by CD68, C1QC, FCGR3A, while EMT shows downregulation of POSTN, COL6A1, COL6A2.These gene-level results are consistent with the program-level claims reported in the paper and support the biological interpretation.
Neighborhood modeling choices: The choice of k=20 and exponential weighting follows standard practice in spatial transcriptomics (neighborhood size of 20–30 cells, distance‑decayed weighting). For larger neighborhoods, the CDS would naturally decrease because local signals are diluted; for alternative weighting schemes, the relative ranking of neighbors remains similar under monotonic transformations. Therefore, the findings are not sensitive to these modeling choices. 
Regarding COMMOT: Optimal transport‑based methods are designed to screen ligand–receptor pairs using predefined interaction databases. Our framework, by contrast, quantifies directional influence without relying on any ligand–receptor priors. 
Counterfactual replacement procedure: Distances are never altered, only sender‑type neighbor identities are replaced using predefined distance bins (e.g., [0,10….40,∞] µm). Replacement cells are chosen from the same core and the same distance bin (with ±1 bin expansion if needed); if none exists, we fall back to any eligible cell (different type for type‑swap, same type for within‑type) in the core.
R #2:
Gene‑set limitation: We agree that program projection depends on predefined gene sets, but CDS significance does not rely on this. Because CDS significance is computed directly from the L1 displacement of residualized cell states in PCA space, independent of any predefined gene sets or functional programs. 
Limited test cores (only 10): Larger multi‑patient cohorts would better capture inter‑patient heterogeneity (noted as a limitation).
Single cancer type: The framework is technology‑agnostic; validation on other cancers is a plan of future work.
Reproducibility (code/data): Source code and data will be made publicly available upon acceptance.
MSE discrepancy: The MSE values in Figure 1 are correct; the text will be corrected accordingly.
R#3:
Causal claims: We fully agree. We will reframe the method as interventional model interpretability (not causal inference) and modify all causal language. Claims 1–2 will be revised accordingly. Assumptions are explicitly listed [spatial locality, residualized state separability, counterfactual stability, model identifiability].
Design choices:
Neighbourhood‑conditioned graph: Receiver state is shaped by local microenvironment, so we model spatial context explicitly. This is a standard choice for counterfactual replacement in spatial data.
Residualization: Subtracting cell‑type mean ensures CDS measures intra‑type functional shifts rather than cell‑type identity differences.
Two swap operators: Type‑swap tests whether sender identity matters, while within‑type shuffle tests whether sender functional state matters. So, together they disentangle identity from heterogeneity.
Comparison with existing methods: Direct numerical comparison with CellChat/NicheNet/GITIII is not feasible because these methods produce fundamentally different outputs (LR co-expression scores and attention weights). In contrast, CDS quantifies a displacement in PCA space after sender perturbation. Thus, these methods are not comparable. However, the manuscript already includes biological validation: CDS correlates strongly with CellChatDB LR scores (Spearman r=0.758, p=0.0027), confirming alignment with known biology. 
Minor corrections: MSE values in Fig 1 will be corrected and duplicate DOIs will be fixed.&lt;/p&gt;

&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Invite for Rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes a counterfactual framework for analyzing directional cell–cell interactions in spatial transcriptomics, with a well-motivated formulation and strong empirical validation, including rigorous null models and biological concordance. Reviewers find the approach promising and methodologically sound, but raise concerns about the interpretation of the method as causal inference, the lack of comparisons to existing approaches, and limited evaluation scope. I suggest a rebuttal to address these concerns.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-2&quot;&gt;Meta-review #2&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The reviewers are overall positive after rebuttal and recommend acceptance. The rebuttal adequately addressed the main concerns, particularly regarding claim framing, methodological clarification, and biological interpretation. Remaining issues mainly concern reproducibility, code/data release, and camera-ready presentation, which should be addressed in the final version. I therefore recommend acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-3&quot;&gt;Meta-review #3&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;After considering the reviews and rebuttal, I recommend acceptance. The paper proposes a counterfactual perturbation framework for analyzing directional cell-cell influence in spatial transcriptomics, with matched null models, FDR control, bootstrap uncertainty estimation, and biological concordance with ligand-receptor knowledge.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Anzum, Humaira AND Kochat, Veena AND Mahmud, Md Ishtyaq AND Dwarampudi, Jagan Mohan Reddy AND Satpati, Suresh AND Shukla, Pooja AND Javle, Milind AND Kwong, Lawrence AND Rai, Kunal AND Banerjee, Tania</name></author><category term="Body -&gt; Abdomen" /><category term="Modalities -&gt; Microscopy" /><category term="Applications -&gt; Computational / Integrative Pathology" /><category term="Applications -&gt; Connectivity &amp; Network Analysis" /><category term="Machine Learning -&gt; Causal Inference / Counterfactual Reasoning" /><category term="Machine Learning -&gt; Deep Learning" /><category term="Machine Learning -&gt; Evaluation / Benchmarking / Reproducibility" /><category term="Machine Learning -&gt; Interpretability / Explainability" /><category term="Machine Learning -&gt; Pathomics" /><category term="Anzum, Humaira" /><category term="Kochat, Veena" /><category term="Mahmud, Md Ishtyaq" /><category term="Dwarampudi, Jagan Mohan Reddy" /><category term="Satpati, Suresh" /><category term="Shukla, Pooja" /><category term="Javle, Milind" /><category term="Kwong, Lawrence" /><category term="Rai, Kunal" /><category term="Banerjee, Tania" /><summary type="html">Abstract Understanding neighboring cell influences is central to spatial transcriptomics, yet existing methods rely on correlation or predefined ligand–receptor (LR) pairs and do not test directionality. We introduce a counterfactual, intervention-based framework for inferring directional cell–cell influence that is LR-agnostic and tests sender specificity. A neighborhood graph model predicts receiver cell state from local spatial context. Directional influence is quantified by counterfactually replacing neighbors of a sender type and measuring resulting displacement in predicted receiver state. We define a Counterfactual Directionality Score (CDS) that quantifies directional influence, and compute pairwise CDS by aggregating across receiver cells and test cores sender–receiver pair. Applied to Xenium cholangiocarcinoma tissue microarrays (38 cores), the framework identified reproducible, asymmetric interactions between tumor, immune, and stromal compartments; e.g. Tumor-EMT →Macrophage (CDS: 0.0828) and Fibroblast→Macrophage (CDS: 0.0582). Effects exceeded label-permutation and spatial-shuffle null models (p &amp;lt; 0.001, FDR-controlled) and remained stable under core-level bootstrap resampling. Inferred directional strengths correlated strongly with matched LR scores (r = 0.758, p = 0.0027), supporting biological concordance. Code:https://github.com/Snigdha022/Counterfactual-Cell-Influence-Spatial-Transcriptomics.git. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/6105_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository https://github.com/Snigdha022/Counterfactual-Cell-Influence-Spatial-Transcriptomics.git Link to the Dataset(s) We utilized spatial transcriptomics data from human cholangiocarcinoma TMAs, selecting 38 independent cores generated via the Xenium platform. These data were obtained from The University of Texas MD Anderson Cancer Center and are not publicly available at this time. The code is demonstrated using a publicly available breast cancer dataset with the following link: https://github.com/Snigdha022/Counterfactual-Cell-Influence-Spatial-Transcriptomics/tree/main/data/raw BibTex @InProceedings{AnzHum_ACounterfactual_MICCAI2026,         author = { Anzum, Humaira AND Kochat, Veena AND Mahmud, Md Ishtyaq AND Dwarampudi, Jagan Mohan Reddy AND Satpati, Suresh AND Shukla, Pooja AND Javle, Milind AND Kwong, Lawrence AND Rai, Kunal AND Banerjee, Tania},         title = { { A Counterfactual Framework for Directional Cell–Cell Interaction Analysis in Spatial Transcriptomics } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16891},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper This paper proposes a novel counterfactual, intervention-based framework for inferring directional cell–cell interactions in spatial transcriptomics. Unlike conventional approaches that rely on correlation or predefined ligand–receptor (LR) priors, the method introduces the Counterfactual Directionality Score (CDS), which quantifies directional influence by explicitly perturbing the sender cell population and measuring the resulting change in the predicted receiver cell state. This formulation enables direct testing of causal directionality in a data-driven manner. The framework is further supported by a neighborhood-conditioned predictive model, rigorous statistical validation using multiple null models with FDR control, and biological validation via concordance with known LR interactions. Overall, the work presents a principled and scalable approach for directional cell–cell communication analysis. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.Methodological novelty through counterfactual inference The proposed framework represents a clear departure from conventional correlation-based or LR-dependent approaches. By introducing explicit counterfactual interventions (e. g. , sender type swapping) and defining CDS, the method directly tests directional influence rather than inferring it indirectly. This is a conceptually significant advancement in spatial transcriptomics analysis. 2.Strong statistical rigor and experimental design The study employs a carefully structured evaluation pipeline, including strict separation of TMA cores into training, validation, and test sets. The use of multiple complementary null models (label permutation, spatial shuffling, and sender-agnostic replacement) with FDR correction ensures robustness against spurious findings. Additionally, core-level bootstrap resampling provides reliable uncertainty quantification. 3.Biological relevance and interpretability Despite being fully data-driven, the inferred interactions show strong concordance with established biological knowledge, as evidenced by a high correlation with LR-based scores. The projection of gene-level changes onto functional programs further enhances interpretability and provides meaningful biological insights that are valuable for both basic and translational research. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.Limited granularity in program-level interpretation While the program-level projection provides a useful summary, the analysis lacks detailed identification of the genes driving each functional program shift. As a result, the biological interpretation remains somewhat high-level and less mechanistically grounded. Suggestion: Reporting top contributing genes and their associated pathways for each program-level effect would significantly improve interpretability. 2.Insufficient analysis of neighborhood modeling choices The use of a fixed number of nearest neighbors (k=20) and a specific exponential distance weighting function is not sufficiently justified or analyzed. The sensitivity of the results to these design choices remains unclear. Suggestion: Performing sensitivity analyses over different values of k and alternative weighting schemes would strengthen the robustness claims. Incorporating more biologically grounded approaches (e. g. , optimal transport-based methods such as COMMOT) could further enhance realism. 3.Lack of clarity in counterfactual replacement procedure The description of how spatial structure is preserved during counterfactual replacement (e. g. , “distance bin structure”) is not sufficiently detailed. This limits reproducibility and raises potential concerns regarding implementation. Suggestion: Providing pseudocode or a more explicit algorithmic description would improve transparency and reproducibility. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html This is a well-motivated and carefully executed study that introduces a meaningful conceptual advance in spatial transcriptomics. The counterfactual formulation is both elegant and practically useful, and the experimental validation is rigorous. To further strengthen the work, the authors are encouraged to improve the granularity of biological interpretation and provide additional clarity on key methodological components, particularly the counterfactual replacement procedure and neighborhood modeling assumptions. Addressing these points would enhance both interpretability and reproducibility. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The primary factors supporting this recommendation are the originality of the proposed counterfactual framework and the strong statistical rigor of the experimental design. The work addresses an important limitation in current spatial transcriptomics methodologies and provides a principled solution that enables explicit testing of directional influence. The results are supported by robust validation strategies and demonstrate meaningful biological consistency. While there are some limitations in interpretability detail and modeling choices, these are relatively minor compared to the overall contribution. Therefore, the paper meets the standard for acceptance at MICCAI. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The rebuttal satisfactorily addresses the main methodological and interpretational concerns raised during review. In particular, the authors appropriately agree to reframe the method as interventional model interpretability rather than formal causal inference, and to revise the causal language accordingly. This revised framing is important because it better aligns the contribution with the actual experiments: a controlled counterfactual perturbation framework for analyzing directional effects in the model-predicted receiver-cell state. The authors also clarify the counterfactual replacement procedure, specifying same-core replacement while preserving distance-bin structure, with fallback rules when exact matches are unavailable. This substantially reduces ambiguity in the implementation and improves reproducibility. The additional clarifications on biological interpretation and neighborhood modeling assumptions are also sufficient for the scope of this work. Some limitations remain, including evaluation on a single cancer type, a limited number of held-out test cores, and the lack of direct numerical comparison with methods (e.g., CellChat, NicheNet, or GITIII). However, I do not consider these limitations fatal. The paper already provides meaningful validation through matched null models, FDR control, bootstrap uncertainty estimation, and ligand–receptor concordance, and the authors reasonably explain that CDS produces a displacement-based counterfactual score that is not directly comparable to ligand–receptor co-expression scores or attention-based attribution outputs. But reproducibility remains an important issue for the impact of this work. My accept-side recommendation assumes that the authors will fulfill their rebuttal commitment to publicly release the source code and relevant data or dependencies upon acceptance. Given the methodological nature of the contribution and the remaining implementation details, code release is important for verifying reproducibility, enabling independent use of the framework, and increasing the paper’s practical impact. I also strongly recommend that the authors carefully revise the manuscript presentation in the camera-ready version. In addition to removing causal overclaims, the paper would benefit from a more formal and polished structure, clearer methodological exposition, and correction of formatting issues (e.g., including standardizing the abstract format as a single coherent paragraph and fixing minor reference/figure-text inconsistencies or all the others (including reviewers’ comments) 100% completely). Overall, the rebuttal resolves the main concerns to my satisfaction. The remaining issues are limitations and camera-ready responsibilities rather than grounds for rejection. I therefore keep my score with weak acceptance. Review #2 Please describe the contribution of the paper The authors present a counterfactual framework for inferring directional cell–cell influence in spatial transcriptomics without relying on predefined ligand–receptor priors. The experimental evaluation reveals reproducible and asymmetric directional influences between tumor, immune, and stromal compartments, most prominently Tumor-EMT→Macrophage and Fibroblast→Macrophage. Three complementary null models confirm that the observed effects are unlikely to arise by chance, and a ligand–receptor concordance analysis supports alignment with established biology (r = +0.758, p = 0.0027). The main contribution of this work is the introduction of a statistically rigorous, intervention-based alternative to correlation-driven cell–cell communication analysis, applicable to large imaging transcriptomics datasets and clinical samples. Central to this contribution is the Counterfactual Directionality Score (CDS), a metric that quantifies directional influence by counterfactually replacing sender cells within spatial neighborhoods and measuring the resulting displacement in the predicted receiver cell state. Together, these elements constitute a principled and scalable framework for dissecting directional signaling in the tumor microenvironment. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The strong concordance between the inferred counterfactual directionality scores and established ligand–receptor scores (r = +0.758, p = 0.0027), which provides meaningful external biological validation and supports the interpretation that the framework captures genuine signaling relationships rather than statistical artifacts. The adoption of an explicit interventional framework that, unlike conventional correlation-based approaches, enables directional causal testing of cell–cell communication in spatial transcriptomics. Ligand–receptor agnostic design, which eliminates dependence on curated interaction databases that are inherently incomplete and subject to annotation bias, thereby broadening the applicability of the framework to less-characterized biological contexts. The rigorous statistical validation strategy, which combines three complementary null models with FDR correction to confirm that observed effects are unlikely to arise by chance. This is further supported by core-level bootstrap resampling and strict train/test separation, together providing robust evidence for the stability and reproducibility of the reported findings. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. A limitation of the biological interpretation layer is that the projection of counterfactual effects onto functional programs is inherently contingent on the quality and completeness of the predefined gene sets used. The experimental evaluation is the relatively small number of held-out test cores used for counterfactual analysis. With only 10 test cores drawn from a tissue microarray, the evaluation may not adequately capture the extent of inter-patient biological variability inherent to cholangiocarcinoma, a disease known for its considerable tumor microenvironment heterogeneity. A minor but clear limitation of the paper is that the proposed framework is evaluated exclusively on a single cancer type (cholangiocarcinoma) which restricts the extent to which the findings and the method’s performance can be generalized to other tissue types, tumor microenvironments, or disease contexts. This could be part of the on going or future work. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html The text reports the MSE values in panel A of Figure 1 for true-NI (0.849), MLP (1.032), and shuffled-NI (0.964), however, the corresponding values visible in the figure are 0.846, 1.036, and 0.969 respectively, indicating a discrepancy between the reported numbers and the actual figure. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (5) Accept — should be accepted, independent of rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? A statistically rigorous, intervention-based alternative to correlation-driven cell–cell communication analysis, applicable to large imaging transcriptomics datasets and clinical samples. Central to this contribution is the Counterfactual Directionality Score (CDS), a metric that quantifies directional influence by counterfactually replacing sender cells within spatial neighborhoods and measuring the resulting displacement in the predicted receiver cell state. Together, these elements constitute a principled and scalable framework for dissecting directional signaling in the tumor microenvironment. In addition, the adoption of an explicit interventional framework that, unlike conventional correlation-based approaches, enables directional causal testing of cell–cell communication in spatial transcriptomics. Finally, LR agnostic design, which eliminates dependence on curated interaction databases that are inherently incomplete and subject to annotation bias, thereby broadening the applicability of the framework to less-characterized biological contexts. Reviewer confidence Somewhat confident (2) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #3 Please describe the contribution of the paper Summary: The paper proposes a counterfactual framework to infer directional cell-cell influences in spatial transcriptomics data without the use of domain priors. The counterfactual framework involves (i) prediction of receiver cell state by a neighbourhood-conditioned graph model, (ii) measuring the displacement in predicted receiver state after controlled perturbations of sender composition. The main claims of the paper are: Claim 1: Existing approaches to inferring cell-cell influences are correlational and either depend on prior knowledge or do not explicitly test directionality. To bridge this gap, they propose a counterfactual framework. Claim 2: The proposed counterfactual framework enables directional inference through intervention testing rather than association. Claim 3: Counterfactual Directionality score (CDS) disentangles cell type identity from intra-type heterogeneity. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The paper proposes an interesting approach for model interpretability rather than causal analysis. The proposed “CDS” score also seems to provide fine-grained interpretability information regarding cell identity and intra-type heterogeneity (Claim 3), as validated in their robustness analysis. It is also good to see the biological interpretation of the model artefacts. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.The paper is jargon-heavy and hard to follow. Revision of some text and providing diagrams (For example: a block diagram for the methodology) would improve the readability of the paper. The motivation for the specific design choices of the methodology was missing. To name a few examples: Why was neighbourhood-conditioned graph model chosen? Has it been applied to similar problem. Why do we need to isolate functional state from cell type identity? Why were the two swap operators selected? It is also necessary to provide evidence from literature for these design choices. 2.Here, I am assessing claim 2 from the methodological perspective, i.e., the question - does the proposed framework enable directional inference of cell-cell communication within the data. To be able to infer the true directional cell-cell influences, the graph model adopted should be close to 100% accurate. Also, one of the main assumptions here is that there is no hidden confounding (measurement or environment induced artefacts). So, what is actually being achieved is model interpretability. Does node A influence node B for my model? If so, what is the direction of influence. The experiments are all going in this direction. 3.Following point 1, claim 1 is invalidated. The proposed methodology is still giving us predictive attribution and association. 4.Experiments comparing the proposed method with the existing methods mentioned in the introduction are missing. As, the proposed methodology is also giving some form of predictive attribution, it is important to assess how the framework performs in comparison to existing methods like CellChat, NicheNet, GITIII,… Please rate the clarity and organization of this paper Poor Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html Detailed constructive comments/action points: 1.The main premise of the paper and the claims need to be reformulated. It will be interesting to sell this as a model interpretability method rather than providing causal guarantees. 2.Comparison to existing methods are necessary. 3.Rephrase the paper extensively for better readability and provide substantial evidence from literature. Also, clearly state all assumptions for the methodology. Minor comment: The doi link seem to appear twice for every reference in the paper. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (2) Reject — should be rejected, independent of rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The fundamental premise of the paper (claiming causal directional inference of cell-cell communications) is methodologically flawed. Even if the authors tweak the claims to frame this as a model interpretability tool, the manuscript still requires substantial revisions to clarify the methodology, as well as the addition of necessary SOTA baseline comparisons (e.g., CellChat, NicheNet) to prove its empirical value. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The authors agreed to modify the claims and state the assumptions explicitly. I am not convinced by the authors’ justification for lack of baseline comparisons, because direct numerical comparison is not required to answer the question “What novel cell-cell influences does the current method discover that the existing approaches miss?”. Despite this, I am recommending acceptance because of the biological perspectives offered in the paper, which is often missing from these kind of interdisciplinary computational papers. Author Feedback We thank all reviewers for their constructive feedback. R#1: Gene‑level interpretability: We agree that reporting the specific genes driving each program improves interpretability. We identified the top three up- and down-regulated genes per program for each significant sender–receiver pair using the same model and data. For example, in Fibroblast→Macrophage, Macrophage_Activation is driven by CD68, C1QC, FCGR3A, while EMT shows downregulation of POSTN, COL6A1, COL6A2.These gene-level results are consistent with the program-level claims reported in the paper and support the biological interpretation. Neighborhood modeling choices: The choice of k=20 and exponential weighting follows standard practice in spatial transcriptomics (neighborhood size of 20–30 cells, distance‑decayed weighting). For larger neighborhoods, the CDS would naturally decrease because local signals are diluted; for alternative weighting schemes, the relative ranking of neighbors remains similar under monotonic transformations. Therefore, the findings are not sensitive to these modeling choices. Regarding COMMOT: Optimal transport‑based methods are designed to screen ligand–receptor pairs using predefined interaction databases. Our framework, by contrast, quantifies directional influence without relying on any ligand–receptor priors. Counterfactual replacement procedure: Distances are never altered, only sender‑type neighbor identities are replaced using predefined distance bins (e.g., [0,10….40,∞] µm). Replacement cells are chosen from the same core and the same distance bin (with ±1 bin expansion if needed); if none exists, we fall back to any eligible cell (different type for type‑swap, same type for within‑type) in the core. R #2: Gene‑set limitation: We agree that program projection depends on predefined gene sets, but CDS significance does not rely on this. Because CDS significance is computed directly from the L1 displacement of residualized cell states in PCA space, independent of any predefined gene sets or functional programs. Limited test cores (only 10): Larger multi‑patient cohorts would better capture inter‑patient heterogeneity (noted as a limitation). Single cancer type: The framework is technology‑agnostic; validation on other cancers is a plan of future work. Reproducibility (code/data): Source code and data will be made publicly available upon acceptance. MSE discrepancy: The MSE values in Figure 1 are correct; the text will be corrected accordingly. R#3: Causal claims: We fully agree. We will reframe the method as interventional model interpretability (not causal inference) and modify all causal language. Claims 1–2 will be revised accordingly. Assumptions are explicitly listed [spatial locality, residualized state separability, counterfactual stability, model identifiability]. Design choices: Neighbourhood‑conditioned graph: Receiver state is shaped by local microenvironment, so we model spatial context explicitly. This is a standard choice for counterfactual replacement in spatial data. Residualization: Subtracting cell‑type mean ensures CDS measures intra‑type functional shifts rather than cell‑type identity differences. Two swap operators: Type‑swap tests whether sender identity matters, while within‑type shuffle tests whether sender functional state matters. So, together they disentangle identity from heterogeneity. Comparison with existing methods: Direct numerical comparison with CellChat/NicheNet/GITIII is not feasible because these methods produce fundamentally different outputs (LR co-expression scores and attention weights). In contrast, CDS quantifies a displacement in PCA space after sender perturbation. Thus, these methods are not comparable. However, the manuscript already includes biological validation: CDS correlates strongly with CellChatDB LR scores (Spearman r=0.758, p=0.0027), confirming alignment with known biology. Minor corrections: MSE values in Fig 1 will be corrected and duplicate DOIs will be fixed. Meta-Review Meta-review #1 Your recommendation Invite for Rebuttal Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. This paper proposes a counterfactual framework for analyzing directional cell–cell interactions in spatial transcriptomics, with a well-motivated formulation and strong empirical validation, including rigorous null models and biological concordance. Reviewers find the approach promising and methodologically sound, but raise concerns about the interpretation of the method as causal inference, the lack of comparisons to existing approaches, and limited evaluation scope. I suggest a rebuttal to address these concerns. After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. N/A Please justify your recommendation. N/A Meta-review #2 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. The reviewers are overall positive after rebuttal and recommend acceptance. The rebuttal adequately addressed the main concerns, particularly regarding claim framing, methodological clarification, and biological interpretation. Remaining issues mainly concern reproducibility, code/data release, and camera-ready presentation, which should be addressed in the final version. I therefore recommend acceptance. Meta-review #3 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. After considering the reviews and rebuttal, I recommend acceptance. The paper proposes a counterfactual perturbation framework for analyzing directional cell-cell influence in spatial transcriptomics, with matched null models, FDR control, bootstrap uncertainty estimation, and biological concordance with ligand-receptor knowledge. back to top</summary></entry><entry><title type="html">A Deep Learning Surrogate Model for Microwave Thermal Ablation: From Preoperative Planning to Intraoperative Re-planning</title><link href="https://papers.miccai.org/miccai-2026/0007-Paper4859" rel="alternate" type="text/html" title="A Deep Learning Surrogate Model for Microwave Thermal Ablation: From Preoperative Planning to Intraoperative Re-planning" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0007-Paper4859</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0007-Paper4859">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;With the increasing interest in thermal ablation therapies, patient-specific predictive simulations have become crucial for improving both preoperative planning and intraoperative corrections. Planning strategies typically rely on the solution of an inverse problem, yet high-fidelity finite element simulations remain computationally prohibitive when embedded within an optimization loop under strict clinical time constraints. In this work, we introduce a real-time deep learning surrogate model for predicting liver microwave ablation outcomes, trained exclusively on synthetic data and personalized through 3D perfusion maps that encode patient anatomy and local perfusion rates. This surrogate model takes as input the antenna position, power and ablation duration, and outputs the predicted necrosis volume within a few milliseconds.
Once integrated into an optimization loop, it becomes possible to propose antenna trajectories and ablation settings that maximize tumor coverage while minimizing damage to healthy tissue. We demonstrate this on clinical data via a physics-based optimization framework that computes optimal ablation parameters in under 2 minutes, critical for preoperative planning or intraoperative re-planning in microwave ablation. The surrogate model was validated on in vivo pre-clinical data where it achieved an average Dice coefficient of 0.82 for necrosis mask prediction.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/4859_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/4859_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;3D-IRCADb-01 dataset: &lt;a href=&quot;https://www.ircad.fr/research/data-sets/liver-segmentation-3d-ircadb-01/&quot;&gt;https://www.ircad.fr/research/data-sets/liver-segmentation-3d-ircadb-01/&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{NahIli_ADeep_MICCAI2026,
        author = { Nahmed, Ilias AND Dettori, Francesco AND Duprez, Michel AND Alvarez, Pablo AND Cotin, Stéphane},
        title = { { A Deep Learning Surrogate Model for Microwave Thermal Ablation: From Preoperative Planning to Intraoperative Re-planning } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16893},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes a deep learning-based surrogate model for fast and patient-specific prediction of thermal damage in microwave ablation (MWA). The method replaces computationally expensive FEM simulations with a 3D neural network that takes as input anatomical perfusion maps, antenna configuration, and treatment parameters, and directly predicts the resulting ablation zone.&lt;/p&gt;

      &lt;p&gt;The key contribution lies in enabling accurate, near real-time estimation of thermal damage, achieving orders-of-magnitude speedup over conventional simulation while maintaining high agreement with both synthetic and in vivo data. Furthermore, the proposed model is integrated into an optimization framework for treatment planning, demonstrating improved tumor coverage and reduced damage to healthy tissue.&lt;/p&gt;

      &lt;p&gt;Overall, the work presents a practical and efficient approach for data-driven MWA planning, bridging the gap between physics-based accuracy and clinical usability.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The proposed deep learning-based surrogate model successfully replaces expensive FEM simulations while maintaining strong predictive accuracy. The combination of a 3D U-Net backbone with conditioning (via FiLM) and attention mechanisms allows the model to capture both local and global dependencies. Notably, the method achieves orders-of-magnitude speedup (milliseconds vs. minutes), which is a significant technical and practical contribution enabling real-time applications.&lt;/p&gt;

      &lt;p&gt;The paper provides a thorough evaluation across synthetic data, patient-specific anatomies, and in vivo experiments, demonstrating robustness and generalization. Beyond prediction accuracy, the authors validate the model in a clinically meaningful downstream task—treatment planning optimization—showing clear improvements in tumor coverage and healthy tissue preservation. This end-to-end demonstration strengthens the practical impact of the work.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.While the proposed surrogate model effectively approximates FEM simulations, it does not explicitly account for temperature-dependent changes in tissue properties (e. g. , thermal conductivity, perfusion, and dielectric properties), which are known to significantly influence heat propagation and ablation outcomes in MWA. This simplification may limit the physical fidelity of the predictions, especially under high-temperature regimes where tissue properties can change nonlinearly. Incorporating or explicitly modeling such effects could further improve the reliability and clinical relevance of the approach.&lt;/p&gt;

      &lt;p&gt;2.The model relies primarily on perfusion maps to represent patient-specific variability, which implicitly captures vascular effects but may not fully account for other important tissue-specific properties. In particular, differences between tumor and healthy tissue in terms of thermal and dielectric properties are not explicitly modeled. This simplification raises concerns about whether perfusion alone is sufficient to capture the full range of factors influencing heat propagation and ablation outcomes.&lt;/p&gt;

      &lt;p&gt;3.The experiments consider treatment durations in the range of 180–500 seconds, which may correspond to regimes where the thermal field is approaching a quasi steady-state. As a result, the model may not sufficiently capture the transient heating dynamics that occur in earlier stages (e. g. , 0–180 seconds), where rapid temperature changes and nonlinear effects are more pronounced. This limited temporal coverage raises concerns about the model’s ability to generalize to shorter-duration treatments or to accurately support real-time intraoperative decision-making during the initial phase of ablation.&lt;/p&gt;

      &lt;p&gt;4.The paper lacks sufficient methodological and implementation details to enable reproducibility. Key aspects such as network architecture specifications, training procedures, data generation pipeline, and hyperparameter settings are not described in enough detail to allow faithful reimplementation. Furthermore, no code or pretrained models are made available. Providing more comprehensive technical details or releasing code would significantly improve transparency and facilitate adoption by the community.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper addresses an important and clinically relevant problem by proposing a fast surrogate model for microwave ablation planning, significantly reducing computation time compared to FEM-based approaches while maintaining strong predictive performance. The work is well-motivated and demonstrates promising results across synthetic, patient-specific, and in vivo settings. In particular, the ability to enable near real-time prediction and integrate with a treatment optimization framework is a notable strength with clear practical implications.&lt;/p&gt;

      &lt;p&gt;However, several limitations prevent a stronger recommendation. From a methodological perspective, the approach relies on relatively standard architectures and simplifies the underlying physics, for example by not explicitly modeling temperature-dependent tissue properties and by primarily relying on perfusion to represent tissue heterogeneity. In addition, the training and evaluation are heavily based on synthetic data, with limited validation on real-world cases, raising concerns about generalization. The restricted range of treatment durations may also limit the model’s ability to capture early transient dynamics. Finally, the paper lacks sufficient implementation details to ensure reproducibility.&lt;/p&gt;

      &lt;p&gt;Overall, while the work is not without weaknesses, it presents a solid and practically meaningful contribution that is likely to be of interest to the MICCAI community, and thus merits a weak accept.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This manuscript presents a surrogate-assisted treatment planning framework for microwave ablation (MWA). The authors use Maxwell + Pennes + Arrhenius FEM simulations to train a 3D U-Net-based surrogate with FiLM and self-attention, conditioned on a 3D perfusion map, antenna location heatmap, power, and duration, to predict the necrotic zone. The surrogate is then embedded into an optimization pipeline for two clinically relevant tasks: preoperative planning and intraoperative replanning after antenna displacement without reinsertion. The study is evaluated on synthetic FEM data, FEM simulations on patient anatomies, and three in vivo porcine liver experiments. The reported speedup over FEM is substantial, and the proposed pipeline addresses a clinically meaningful problem.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The clinical motivation is well defined. Rather than aiming broadly at faster simulation, the work targets two concrete bottlenecks in current MWA workflows: treatment planning before the procedure and rapid replanning after antenna misplacement. 
2.The engineering pipeline is relatively complete, since the surrogate is not presented only as a forward predictor but is integrated into a usable planning and replanning framework. 
3.The focus on perfusion and heat-sink effects is well justified by the supporting physical analysis, and this is an appropriate modeling direction. 
4.It is noteworthy that the model is trained entirely on synthetic data yet still shows reasonable performance on patient anatomies and in vivo porcine experiments.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Although perfusion and heat-sink effects are presented as central to the method, the experimental results are reported mostly as global averages. The paper lacks stratified analysis on the difficult cases where these factors matter most, such as tumors close to major vessels, regions with strong local cooling, larger tumors, or geometrically elongated tumors. This issue affects both the necrosis prediction and the planning results. The authors already have the ingredients for such an analysis, including vessel segmentation, radius estimation, perfusion maps, tumor volume, and tumor shape. Reporting performance by vascular proximity, perfusion intensity, tumor size, and shape complexity would make the claims much more convincing and would better define the boundary of applicability of the method.
2.The paper emphasizes patient-specific perfusion, but it does not directly test how robust the surrogate is to physiological distribution shift. The synthetic training distribution appears relatively constrained, whereas the supplementary material suggests substantial inter-subject variability in perfusion. As a result, the current results demonstrate anatomy generalization more clearly than physiological generalization. A dedicated robustness analysis across in-distribution, boundary, and mildly out-of-distribution perfusion conditions would be important, and this should be feasible using synthetic data alone.
3.The validation of the optimization stage is still incomplete. For the 112 clinical planning cases, the main comparison is against the manufacturer chart, but this does not establish whether surrogate-guided optimization is consistent with high-fidelity FEM optimization. Since the planning stage depends on the surrogate not only for forward prediction but also for ranking candidate solutions, the paper should include FEM back-checking on at least a representative subset of optimized cases. Comparing surrogate-selected solutions with FEM-evaluated solutions, or ideally with FEM-optimal solutions, would provide much stronger support for the optimization claims.
4.The replanning experiments are too restricted to support a broad claim of intraoperative robustness. The current setup considers only a 5 mm axial displacement and optimizes a limited subset of parameters. This demonstrates feasibility in one simplified scenario, but it is not sufficient to show robustness to the broader range of intraoperative deviations that may occur in practice. At minimum, the authors should include sensitivity experiments with different displacement magnitudes, non-axial translations, and small angular perturbations, even if these are first performed in synthetic or FEM-based settings.
5.Some of the paper’s stronger claims are not yet sufficiently supported by targeted evidence. In particular, the contribution of FiLM and self-attention remains somewhat unclear because the ablation results appear close in overall performance, and the practical value of patient-specific perfusion is not fully disentangled from the gain brought simply by using an optimizer stronger than the manufacturer chart. These issues could be addressed by more targeted comparisons: evaluating architectural ablations on difficult subsets, and comparing planning performance under homogeneous perfusion, image-derived vessel perfusion, and fully personalized perfusion within the same optimization framework.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;I recommend “weak accept” for this work. It is noteworthy that the model, trained entirely on synthetic data, still demonstrates reasonable performance on patient anatomies and in vivo porcine experiments. The reported speedup over FEM is substantial, and the proposed pipeline effectively addresses a clinically relevant challenge.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors have developed a deep learning surrogate model for real-time microwave thermal ablation outcome prediction, which can be used for both pre- and intra-operative planning. It combines a 3D U-Net architecture with FiLM conditioning and self-attention to encode 3D perfusion maps and ablation parameters.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The proposed model decreases significantly the simulation/prediction time of traditional numerical approaches, e.g., FEM, from minutes to milliseconds, which makes it a strong tool especially for intra-operative replanning where real-time performance is essential. Another strength of the paper is that the model is trained exclusively on simulated data. This effectively tackles the common problem of data insufficiency in the medical field, demonstrating a viable path forward when clinical data is limited.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The specific target tissues are not mentioned in the Abstract or the main body until Section 2.3, where hepatic vascular trees are finally introduced. This key information should be presented much earlier to provide necessary context.
2.While the authors mention that the “The model is conditioned on both intervention- and patient specific parameters”, then, they mention that “once trained, our network can be deployed without per-case retraining” and “we promote generalization across patients and clinical scenarios without requiring per-case re training”. This indicates that, in fact, the patient-specific factor is not taken into account during inference. 
3.The paper lacks a comparison to other existing studies or state-of-the-art methods. As a result, it is unclear if the achieved Dice score is clinically sufficient; in this context, leaving tumor tissue or removing excess healthy tissue would have severe adverse effects on the patient. The authors should address this. 
4.While the validation effort is noted, it is insufficient as there is no clinical ground truth provided.
5.The authors mention interpretability as a feature of the work, but they do not provide any examples or visualizations of interpretable results to support this claim.
7.The statement “Each vessel segment was dilated to radii sampled from a normal distribution based on phsyiological literature” needs to be supported by a reference.
8.Typos and grammar errors:
-The phrase “after antenna misplacement approximately 30 s” is unclear; it is uncertain if the authors mean “within” a 30-s window.
-The text uses both “FEM” and “FE” interchangeably; a single abbreviation should be used. Also, “MW ablations” should become “MWAs”
-Several typos were noted, such as “phsyiological” and the repetition in “network can be be deployed.”&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(5) Accept — should be accepted, independent of rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Traditional numerical methods (FEM, CFD, etc.) often suffer from limited computational efficiency. This bottleneck, combined with the general lack of clinical validation data, makes the integration of computational modeling into the medical field a difficult task. The major factor for my recommendation is that this paper has effectively “eliminated” the computational time barrier by providing results comparable to FEM in a fraction of the time. Furthermore, the authors have achieved a meaningful level of validation despite data scarcity challenges. While there are minor points to be addressed regarding clarity and formatting, the core contribution, enabling near-instantaneous simulation, is significant enough to warrant acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;We thank the reviewers for their positive assessment and constructive comments. We appreciate that the clinical motivation, computational speed-up, and validation effort were recognized. Due to space limits, we grouped the comments below:
Choice of perfusion (R1,R2):
We focused on perfusion because the bioheat equation is particularly sensitive to this parameter, as shown in our previous work. The input is a voxelwise perfusion map encoding anatomy, vessels and lesions. Depending on tumor type, perfusion may be higher or lower than that of healthy liver tissue. Here, tumors were modeled as hypervascular, but the same synthetic framework can accommodate other tumor types by adapting the sampled ranges. Thermal and dielectric properties are relevant, but this work deliberately focused on perfusion as the main variable.
Patient-specific modeling (R3):
By “patient-specific”, we refer to the input conditioning: the network is trained once, while inference uses the patient anatomy and perfusion map. Thus, the model is general in its weights but patient-specific in its inputs.
Ablation-duration and replanning-perturbation ranges (R1,R2):
The 180–500 s range was chosen to match clinical practice. Manufacturer charts are also provided in this range. The framework itself is not restricted to this interval, since synthetic training can include shorter or longer durations. Regarding replanning, validation used a 5 mm axial perturbation to remain consistent with our previous FE-based study. The optimization itself used a wider axial translation range and a stochastic gradient-free method. Non-axial translations and angular perturbations are important for planning validation, but fall outside the scope of replanning, which is deliberately constrained to corrections along the existing insertion axis to maximize tumor coverage while avoiding reinsertion.
Synthetic data analysis, architecture, and interpretability (R2,R3):
We acknowledge that physiological distribution was not exhaustively investigated. However, the in vivo ablations were performed under different local perfusion conditions, providing partial verification of robustness. A deeper analysis would provide stronger evidence, but within space constraints, we prioritized this external validation. Likewise, a full network ablation study with more cases would better highlight the architecture. However, we could only fit the compact study in Table 1, where the proposed architecture improved Dice on in vivo experiments by 2 points on average compared to the baselines. Finally, “interpretability” referred to the planning framework vis-a-vis our concurrent submission. As mentioned in the conclusion, FE simulations can still be integrated near the end of the loop to verify that the proposed plan remains physically consistent.
Planning validation (R2):
In the present work, the consistency of the planning tool was partially assessed through the replanning validation: we reproduced the setup of our previous FE-based replanning study and obtained similar results with a faster approach.
Comparison, Dice, and healthy tissue damage (R3):
To our knowledge, there are no directly comparable NN-based MWA planning studies. Existing surrogate studies mainly concern RFA, where the modality and control variables differ, making direct comparison not quite fair. Numerical MWA planning studies report Dice scores in a similar range [4,7,8,9]. To assess healthy tissue damage, we reported both coverage and Dice: coverage prioritizes complete tumor treatment, while Dice penalizes excessive ablation outside the target. The average Dice improved from 0.62 to 0.80, suggesting reduced unnecessary damage.
Minor revisions (R3):
In the camera-ready submission, we moved the target tissue information earlier, added the requested reference for vessel-radius sampling, and corrected the reported typos and terminology inconsistencies.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Provisional Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Overall all reviewers found the method novel and interesting.&lt;/p&gt;

      &lt;p&gt;However, there were several points of clarify reviewers (especially #1 and #2) ask for that could be addressed to strengthen scores. These include the range of parameters considered during training, how the model was validated, and comparison methods.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Nahmed, Ilias AND Dettori, Francesco AND Duprez, Michel AND Alvarez, Pablo AND Cotin, Stéphane</name></author><category term="Body -&gt; Abdomen" /><category term="Modalities -&gt; CT / X-ray" /><category term="Applications -&gt; Digital Twins / Simulation / Synthetic Data" /><category term="Applications -&gt; Outcome Prediction / Prognosis / Longitudinal Modeling" /><category term="Machine Learning -&gt; Deep Learning" /><category term="Machine Learning -&gt; Synthetic Data / Data-centric AI" /><category term="Surgery -&gt; Planning &amp; Simulation" /><category term="Nahmed, Ilias" /><category term="Dettori, Francesco" /><category term="Duprez, Michel" /><category term="Alvarez, Pablo" /><category term="Cotin, Stéphane" /><summary type="html">Abstract With the increasing interest in thermal ablation therapies, patient-specific predictive simulations have become crucial for improving both preoperative planning and intraoperative corrections. Planning strategies typically rely on the solution of an inverse problem, yet high-fidelity finite element simulations remain computationally prohibitive when embedded within an optimization loop under strict clinical time constraints. In this work, we introduce a real-time deep learning surrogate model for predicting liver microwave ablation outcomes, trained exclusively on synthetic data and personalized through 3D perfusion maps that encode patient anatomy and local perfusion rates. This surrogate model takes as input the antenna position, power and ablation duration, and outputs the predicted necrosis volume within a few milliseconds. Once integrated into an optimization loop, it becomes possible to propose antenna trajectories and ablation settings that maximize tumor coverage while minimizing damage to healthy tissue. We demonstrate this on clinical data via a physics-based optimization framework that computes optimal ablation parameters in under 2 minutes, critical for preoperative planning or intraoperative re-planning in microwave ablation. The surrogate model was validated on in vivo pre-clinical data where it achieved an average Dice coefficient of 0.82 for necrosis mask prediction. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/4859_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository N/A Link to the Dataset(s) 3D-IRCADb-01 dataset: https://www.ircad.fr/research/data-sets/liver-segmentation-3d-ircadb-01/ BibTex @InProceedings{NahIli_ADeep_MICCAI2026,         author = { Nahmed, Ilias AND Dettori, Francesco AND Duprez, Michel AND Alvarez, Pablo AND Cotin, Stéphane},         title = { { A Deep Learning Surrogate Model for Microwave Thermal Ablation: From Preoperative Planning to Intraoperative Re-planning } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16893},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper This paper proposes a deep learning-based surrogate model for fast and patient-specific prediction of thermal damage in microwave ablation (MWA). The method replaces computationally expensive FEM simulations with a 3D neural network that takes as input anatomical perfusion maps, antenna configuration, and treatment parameters, and directly predicts the resulting ablation zone. The key contribution lies in enabling accurate, near real-time estimation of thermal damage, achieving orders-of-magnitude speedup over conventional simulation while maintaining high agreement with both synthetic and in vivo data. Furthermore, the proposed model is integrated into an optimization framework for treatment planning, demonstrating improved tumor coverage and reduced damage to healthy tissue. Overall, the work presents a practical and efficient approach for data-driven MWA planning, bridging the gap between physics-based accuracy and clinical usability. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The proposed deep learning-based surrogate model successfully replaces expensive FEM simulations while maintaining strong predictive accuracy. The combination of a 3D U-Net backbone with conditioning (via FiLM) and attention mechanisms allows the model to capture both local and global dependencies. Notably, the method achieves orders-of-magnitude speedup (milliseconds vs. minutes), which is a significant technical and practical contribution enabling real-time applications. The paper provides a thorough evaluation across synthetic data, patient-specific anatomies, and in vivo experiments, demonstrating robustness and generalization. Beyond prediction accuracy, the authors validate the model in a clinically meaningful downstream task—treatment planning optimization—showing clear improvements in tumor coverage and healthy tissue preservation. This end-to-end demonstration strengthens the practical impact of the work. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.While the proposed surrogate model effectively approximates FEM simulations, it does not explicitly account for temperature-dependent changes in tissue properties (e. g. , thermal conductivity, perfusion, and dielectric properties), which are known to significantly influence heat propagation and ablation outcomes in MWA. This simplification may limit the physical fidelity of the predictions, especially under high-temperature regimes where tissue properties can change nonlinearly. Incorporating or explicitly modeling such effects could further improve the reliability and clinical relevance of the approach. 2.The model relies primarily on perfusion maps to represent patient-specific variability, which implicitly captures vascular effects but may not fully account for other important tissue-specific properties. In particular, differences between tumor and healthy tissue in terms of thermal and dielectric properties are not explicitly modeled. This simplification raises concerns about whether perfusion alone is sufficient to capture the full range of factors influencing heat propagation and ablation outcomes. 3.The experiments consider treatment durations in the range of 180–500 seconds, which may correspond to regimes where the thermal field is approaching a quasi steady-state. As a result, the model may not sufficiently capture the transient heating dynamics that occur in earlier stages (e. g. , 0–180 seconds), where rapid temperature changes and nonlinear effects are more pronounced. This limited temporal coverage raises concerns about the model’s ability to generalize to shorter-duration treatments or to accurately support real-time intraoperative decision-making during the initial phase of ablation. 4.The paper lacks sufficient methodological and implementation details to enable reproducibility. Key aspects such as network architecture specifications, training procedures, data generation pipeline, and hyperparameter settings are not described in enough detail to allow faithful reimplementation. Furthermore, no code or pretrained models are made available. Providing more comprehensive technical details or releasing code would significantly improve transparency and facilitate adoption by the community. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? This paper addresses an important and clinically relevant problem by proposing a fast surrogate model for microwave ablation planning, significantly reducing computation time compared to FEM-based approaches while maintaining strong predictive performance. The work is well-motivated and demonstrates promising results across synthetic, patient-specific, and in vivo settings. In particular, the ability to enable near real-time prediction and integrate with a treatment optimization framework is a notable strength with clear practical implications. However, several limitations prevent a stronger recommendation. From a methodological perspective, the approach relies on relatively standard architectures and simplifies the underlying physics, for example by not explicitly modeling temperature-dependent tissue properties and by primarily relying on perfusion to represent tissue heterogeneity. In addition, the training and evaluation are heavily based on synthetic data, with limited validation on real-world cases, raising concerns about generalization. The restricted range of treatment durations may also limit the model’s ability to capture early transient dynamics. Finally, the paper lacks sufficient implementation details to ensure reproducibility. Overall, while the work is not without weaknesses, it presents a solid and practically meaningful contribution that is likely to be of interest to the MICCAI community, and thus merits a weak accept. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #2 Please describe the contribution of the paper This manuscript presents a surrogate-assisted treatment planning framework for microwave ablation (MWA). The authors use Maxwell + Pennes + Arrhenius FEM simulations to train a 3D U-Net-based surrogate with FiLM and self-attention, conditioned on a 3D perfusion map, antenna location heatmap, power, and duration, to predict the necrotic zone. The surrogate is then embedded into an optimization pipeline for two clinically relevant tasks: preoperative planning and intraoperative replanning after antenna displacement without reinsertion. The study is evaluated on synthetic FEM data, FEM simulations on patient anatomies, and three in vivo porcine liver experiments. The reported speedup over FEM is substantial, and the proposed pipeline addresses a clinically meaningful problem. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The clinical motivation is well defined. Rather than aiming broadly at faster simulation, the work targets two concrete bottlenecks in current MWA workflows: treatment planning before the procedure and rapid replanning after antenna misplacement. 2.The engineering pipeline is relatively complete, since the surrogate is not presented only as a forward predictor but is integrated into a usable planning and replanning framework. 3.The focus on perfusion and heat-sink effects is well justified by the supporting physical analysis, and this is an appropriate modeling direction. 4.It is noteworthy that the model is trained entirely on synthetic data yet still shows reasonable performance on patient anatomies and in vivo porcine experiments. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.Although perfusion and heat-sink effects are presented as central to the method, the experimental results are reported mostly as global averages. The paper lacks stratified analysis on the difficult cases where these factors matter most, such as tumors close to major vessels, regions with strong local cooling, larger tumors, or geometrically elongated tumors. This issue affects both the necrosis prediction and the planning results. The authors already have the ingredients for such an analysis, including vessel segmentation, radius estimation, perfusion maps, tumor volume, and tumor shape. Reporting performance by vascular proximity, perfusion intensity, tumor size, and shape complexity would make the claims much more convincing and would better define the boundary of applicability of the method. 2.The paper emphasizes patient-specific perfusion, but it does not directly test how robust the surrogate is to physiological distribution shift. The synthetic training distribution appears relatively constrained, whereas the supplementary material suggests substantial inter-subject variability in perfusion. As a result, the current results demonstrate anatomy generalization more clearly than physiological generalization. A dedicated robustness analysis across in-distribution, boundary, and mildly out-of-distribution perfusion conditions would be important, and this should be feasible using synthetic data alone. 3.The validation of the optimization stage is still incomplete. For the 112 clinical planning cases, the main comparison is against the manufacturer chart, but this does not establish whether surrogate-guided optimization is consistent with high-fidelity FEM optimization. Since the planning stage depends on the surrogate not only for forward prediction but also for ranking candidate solutions, the paper should include FEM back-checking on at least a representative subset of optimized cases. Comparing surrogate-selected solutions with FEM-evaluated solutions, or ideally with FEM-optimal solutions, would provide much stronger support for the optimization claims. 4.The replanning experiments are too restricted to support a broad claim of intraoperative robustness. The current setup considers only a 5 mm axial displacement and optimizes a limited subset of parameters. This demonstrates feasibility in one simplified scenario, but it is not sufficient to show robustness to the broader range of intraoperative deviations that may occur in practice. At minimum, the authors should include sensitivity experiments with different displacement magnitudes, non-axial translations, and small angular perturbations, even if these are first performed in synthetic or FEM-based settings. 5.Some of the paper’s stronger claims are not yet sufficiently supported by targeted evidence. In particular, the contribution of FiLM and self-attention remains somewhat unclear because the ablation results appear close in overall performance, and the practical value of patient-specific perfusion is not fully disentangled from the gain brought simply by using an optimizer stronger than the manufacturer chart. These issues could be addressed by more targeted comparisons: evaluating architectural ablations on difficult subsets, and comparing planning performance under homogeneous perfusion, image-derived vessel perfusion, and fully personalized perfusion within the same optimization framework. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? I recommend “weak accept” for this work. It is noteworthy that the model, trained entirely on synthetic data, still demonstrates reasonable performance on patient anatomies and in vivo porcine experiments. The reported speedup over FEM is substantial, and the proposed pipeline effectively addresses a clinically relevant challenge. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #3 Please describe the contribution of the paper The authors have developed a deep learning surrogate model for real-time microwave thermal ablation outcome prediction, which can be used for both pre- and intra-operative planning. It combines a 3D U-Net architecture with FiLM conditioning and self-attention to encode 3D perfusion maps and ablation parameters. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The proposed model decreases significantly the simulation/prediction time of traditional numerical approaches, e.g., FEM, from minutes to milliseconds, which makes it a strong tool especially for intra-operative replanning where real-time performance is essential. Another strength of the paper is that the model is trained exclusively on simulated data. This effectively tackles the common problem of data insufficiency in the medical field, demonstrating a viable path forward when clinical data is limited. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.The specific target tissues are not mentioned in the Abstract or the main body until Section 2.3, where hepatic vascular trees are finally introduced. This key information should be presented much earlier to provide necessary context. 2.While the authors mention that the “The model is conditioned on both intervention- and patient specific parameters”, then, they mention that “once trained, our network can be deployed without per-case retraining” and “we promote generalization across patients and clinical scenarios without requiring per-case re training”. This indicates that, in fact, the patient-specific factor is not taken into account during inference. 3.The paper lacks a comparison to other existing studies or state-of-the-art methods. As a result, it is unclear if the achieved Dice score is clinically sufficient; in this context, leaving tumor tissue or removing excess healthy tissue would have severe adverse effects on the patient. The authors should address this. 4.While the validation effort is noted, it is insufficient as there is no clinical ground truth provided. 5.The authors mention interpretability as a feature of the work, but they do not provide any examples or visualizations of interpretable results to support this claim. 7.The statement “Each vessel segment was dilated to radii sampled from a normal distribution based on phsyiological literature” needs to be supported by a reference. 8.Typos and grammar errors: -The phrase “after antenna misplacement approximately 30 s” is unclear; it is uncertain if the authors mean “within” a 30-s window. -The text uses both “FEM” and “FE” interchangeably; a single abbreviation should be used. Also, “MW ablations” should become “MWAs” -Several typos were noted, such as “phsyiological” and the repetition in “network can be be deployed.” Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (5) Accept — should be accepted, independent of rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? Traditional numerical methods (FEM, CFD, etc.) often suffer from limited computational efficiency. This bottleneck, combined with the general lack of clinical validation data, makes the integration of computational modeling into the medical field a difficult task. The major factor for my recommendation is that this paper has effectively “eliminated” the computational time barrier by providing results comparable to FEM in a fraction of the time. Furthermore, the authors have achieved a meaningful level of validation despite data scarcity challenges. While there are minor points to be addressed regarding clarity and formatting, the core contribution, enabling near-instantaneous simulation, is significant enough to warrant acceptance. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Author Feedback We thank the reviewers for their positive assessment and constructive comments. We appreciate that the clinical motivation, computational speed-up, and validation effort were recognized. Due to space limits, we grouped the comments below: Choice of perfusion (R1,R2): We focused on perfusion because the bioheat equation is particularly sensitive to this parameter, as shown in our previous work. The input is a voxelwise perfusion map encoding anatomy, vessels and lesions. Depending on tumor type, perfusion may be higher or lower than that of healthy liver tissue. Here, tumors were modeled as hypervascular, but the same synthetic framework can accommodate other tumor types by adapting the sampled ranges. Thermal and dielectric properties are relevant, but this work deliberately focused on perfusion as the main variable. Patient-specific modeling (R3): By “patient-specific”, we refer to the input conditioning: the network is trained once, while inference uses the patient anatomy and perfusion map. Thus, the model is general in its weights but patient-specific in its inputs. Ablation-duration and replanning-perturbation ranges (R1,R2): The 180–500 s range was chosen to match clinical practice. Manufacturer charts are also provided in this range. The framework itself is not restricted to this interval, since synthetic training can include shorter or longer durations. Regarding replanning, validation used a 5 mm axial perturbation to remain consistent with our previous FE-based study. The optimization itself used a wider axial translation range and a stochastic gradient-free method. Non-axial translations and angular perturbations are important for planning validation, but fall outside the scope of replanning, which is deliberately constrained to corrections along the existing insertion axis to maximize tumor coverage while avoiding reinsertion. Synthetic data analysis, architecture, and interpretability (R2,R3): We acknowledge that physiological distribution was not exhaustively investigated. However, the in vivo ablations were performed under different local perfusion conditions, providing partial verification of robustness. A deeper analysis would provide stronger evidence, but within space constraints, we prioritized this external validation. Likewise, a full network ablation study with more cases would better highlight the architecture. However, we could only fit the compact study in Table 1, where the proposed architecture improved Dice on in vivo experiments by 2 points on average compared to the baselines. Finally, “interpretability” referred to the planning framework vis-a-vis our concurrent submission. As mentioned in the conclusion, FE simulations can still be integrated near the end of the loop to verify that the proposed plan remains physically consistent. Planning validation (R2): In the present work, the consistency of the planning tool was partially assessed through the replanning validation: we reproduced the setup of our previous FE-based replanning study and obtained similar results with a faster approach. Comparison, Dice, and healthy tissue damage (R3): To our knowledge, there are no directly comparable NN-based MWA planning studies. Existing surrogate studies mainly concern RFA, where the modality and control variables differ, making direct comparison not quite fair. Numerical MWA planning studies report Dice scores in a similar range [4,7,8,9]. To assess healthy tissue damage, we reported both coverage and Dice: coverage prioritizes complete tumor treatment, while Dice penalizes excessive ablation outside the target. The average Dice improved from 0.62 to 0.80, suggesting reduced unnecessary damage. Minor revisions (R3): In the camera-ready submission, we moved the target tissue information earlier, added the requested reference for vessel-radius sampling, and corrected the reported typos and terminology inconsistencies. Meta-Review Meta-review #1 Your recommendation Provisional Accept Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. Overall all reviewers found the method novel and interesting. However, there were several points of clarify reviewers (especially #1 and #2) ask for that could be addressed to strengthen scores. These include the range of parameters considered during training, how the model was validated, and comparison methods. back to top</summary></entry><entry><title type="html">A Flexible Structure-Guided Feature Aggregation Paradigm for Efficient Vascular Representation</title><link href="https://papers.miccai.org/miccai-2026/0008-Paper4457" rel="alternate" type="text/html" title="A Flexible Structure-Guided Feature Aggregation Paradigm for Efficient Vascular Representation" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0008-Paper4457</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0008-Paper4457">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;Accurate representation of vascular structures is fundamental for a wide range of clinical applications and remains challenging due to the slender geometry, complex branching topology, and long-range continuity inherent to tree-like vasculature. Although structural priors have been widely explored in vascular image analysis, existing approaches typically integrate such priors implicitly or isotropically during feature aggregation, limiting their ability to faithfully capture directional and structural dependencies along the vessels. In this study, we propose a flexible, structure-guided feature aggregation paradigm that explicitly incorporates anatomical structural cues into the feature aggregation process for efficient vascular representation. The core idea is to aggregate features along directionally guided paths that follow local vascular structures, enabling consistent propagation of structural information while maintaining computational efficiency. To avoid indiscriminate prior enforcement, the proposed paradigm further introduces an adaptive mechanism that selectively injects structural guidance where representation is most challenging. This aggregation paradigm is lightweight, modular, and can be seamlessly integrated into multi-stage feature hierarchies without altering the backbone architecture. Extensive experiments on vascular imaging tasks show that our approach consistently improves the quality of vascular representation. Our code is available at: https://github.com/YaoleiQi/VSP.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/4457_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/4457_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/YaoleiQi/VSP&quot;&gt;https://github.com/YaoleiQi/VSP&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{QiYao_AFlexible_MICCAI2026,
        author = { Qi, Yaolei AND Peng, Wenbo AND Wang, Tong AND Xu, Chenwei AND Zhang, Yuan AND Yang, Guanyu},
        title = { { A Flexible Structure-Guided Feature Aggregation Paradigm for Efficient Vascular Representation } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16896},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The manuscript introduces VSP-Branch, a structure-guided feature aggregation module designed to enhance vascular representation by sampling features along learned anatomical paths and modulating their injection via a Difficulty Gate.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The VSP-Branch demonstrates strong architectural flexibility, functioning as a plug-and-play module that integrates effectively into CNN, ViT, and Mamba-based backbones.&lt;/p&gt;

      &lt;p&gt;2.The method achieves consistent empirical improvements across a diverse range of imaging modalities and scales, encompassing both 2D (DCA1, DeepRetina) and 3D (COSTA, KIPA) datasets.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The path construction mechanism lacks formulation and ablation for 3D topologies. The methodology section explicitly defines offset recursion using a standard 2D rotation matrix. However, the framework is applied to 3D datasets (COSTA, KIPA) without detailing how the degrees of freedom (e. g. , pitch and yaw) are parameterized or constrained in 3D space. Furthermore, the ablation study (Table 2) validates the path sampling steps and step-size ranges solely on the 2D DCA1 dataset, leaving the stability of 3D path generation unverified.&lt;/p&gt;

      &lt;p&gt;2.The functional claim regarding the Difficulty Gate is insufficiently supported by quantitative evidence. The manuscript claims the gate adaptively modulates injection strength in difficult regions based on qualitative visual overlays. Because the gate is supervised only implicitly via the final segmentation objective, there is no quantitative metric to prove it learns to target topological ambiguity rather than defaulting to standard edge detection. 
3.The experimental results lack statistical significance testing. While the VSP-Branch consistently improves metrics across baselines, the absolute gains in metrics such as Hausdorff Distance (HD) are frequently marginal (e. g. , HD improvements of less than 0.1 on DCA1). Without p-values, it is difficult to ascertain whether these improvements reflect genuine structural enhancement or mere statistical variance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors claimed to release the source code and/or dataset upon acceptance of the submission.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This manuscript demonstrates substantial engineering value and a comprehensive empirical workload, despite notable gaps in theoretical formulation. The primary factor driving this score is the high practical utility of the VSP-Branch; it functions as a highly flexible module that achieves consistent performance improvements across diverse baseline architectures, including CNNs, ViTs, and Mamba models, with negligible computational overhead. Furthermore, the authors present a solid validation workload spanning multiple 2D and 3D datasets. However, I withheld a higher score due to specific analytical deficiencies. Most critically, while the method is applied to 3D data, the path generation mechanism is mathematically formulated and ablated strictly within a 2D context using 2D rotation matrices, leaving the 3D topological implementation ambiguous. Additionally, the claim that the Difficulty Gate selectively targets complex anatomical regions relies heavily on qualitative visual overlays rather than rigorous quantitative metrics or statistical significance testing, making it difficult to verify the robustness of marginal metric gains. Overall, the module’s architectural flexibility and strong empirical baseline constitute a net positive contribution to the community that outweighs these presentational shortcomings.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper proposes a plug-and-play module called VSP-Branch to explicitly incorporate anatomical cues into the aggregation process, improving vascular representation efficiency. Extensive experiments show VSP-Branch outperforms other methods on Dice metric.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The figures provided in the manuscript are well-presented and highly clear.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The “Difficulty Gate” is composed of merely a simple convolution. It is unclear how such a straightforward design achieves the authors’ claim of “emphasizing continuity cues more strongly in regions that are difficult to represent while preserving the backbone feature in simpler regions.” Since the Difficulty Gate lacks explicit additional supervision during the training process, how does the module actually learn to distinguish between “simple” and “difficult” regions?
2.Following previous works such as DSCNet, the authors should include additional metrics to evaluate the topological connectivity of the segmented vessels, such as $\beta_0$ and $\beta_1$. 
3.In Fig. 3, bounding boxes should be provided on the original full-size images to indicate the cropped areas for the renal dataset visualizations. Furthermore, the visualization of the Difficulty Gate Map is quite confusing. What do the different colored lines represent? The vessels within the areas enclosed by the blue lines appear very clear and do not seem challenging to segment. The authors should provide the baseline segmentation results (“without VSP-Branch”) for comparison. This would help verify whether the VSP-Branch accurately targets and localizes difficult regions in areas that are otherwise prone to mis-segmentation.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The conceptual design of the structure-guided feature aggregation in the VSP-Branch aligns well with the unique characteristics of the vessel segmentation task. However, regarding the core module—the Difficulty Gate Map—both the methodological design and the visualization results in Fig. 3 fail to convincingly demonstrate its claimed role in reinforcing vessel features specifically in difficult regions.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Reject&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;While the authors provided quantitative improvements, the core concerns remain unresolved:
1.Unconvincing Mechanism: It remains theoretically unjustified how a simple, unsupervised convolution layer can inherently distinguish between simple and difficult regions. Quantitative gains alone cannot validate a module with questionable underlying logic.
2.Contradictory Visualization: Figure 3 fails to show the Difficulty Gate focusing on genuinely challenging areas (e.g., thin, low-contrast vessels). Furthermore, the authors failed to clarify the meaning of the colored lines, leaving the visualization uninterpretable.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper introduces a structure-guided feature aggregation module, termed VSP-Branch, for improving vascular representation. Rather than relying on conventional local convolutional aggregation or incorporating structural priors only through loss functions, the method aims to explicitly model feature interactions along learned, structure-aligned trajectories. Concretely, it predicts per-pixel directions, step sizes, and turning angles to construct sampling paths, along which features are aggregated to capture long-range structural dependencies. A difficulty-aware gating mechanism is further employed to adaptively control the extent of structural information injected at different spatial locations. The module is lightweight and designed to be easily integrated into various backbone architectures, including CNNs, Transformers, and Mamba-based models. Experimental results across several 2D and 3D vascular datasets show consistent improvements over strong baseline methods.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The paper introduces a representation-level structural aggregation mechanism, which differs from most existing approaches that incorporate structural priors either via specialized convolution operators or loss functions. Modeling feature propagation along learned paths is an interesting and non-trivial design. 
2.The proposed VSP-Branch is modular and can be easily plugged into different backbone architectures with negligible computational overhead. This significantly improves its practical applicability. 
3.The method is evaluated across multiple datasets and backbone families (CNN, ViT, Mamba), showing consistent performance gains. 
4.The Difficulty Gate prevents over-enforcement of structural priors and allows the model to focus on challenging regions such as thin vessels and bifurcations. 
5.The paper provides a reasonable interpretation that vascular features exhibit semantic consistency along trajectories, and the proposed method aligns well with this intuition.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ul&gt;
        &lt;li&gt;The path construction is learned solely from the segmentation objective without explicit structural supervision. This raises concerns about the reliability of the learned paths, especially in background or vessel-dense regions, where paths may cross background and unintentionally connect nearby vessels. Such behavior could lead to undesired feature mixing and weaken structural consistency, but this issue is not analyzed in the paper.&lt;/li&gt;
        &lt;li&gt;The method predicts a single path for each pixel, which may be limiting in bifurcation regions where multiple valid directions exist. Since vascular structures are inherently tree-like, it is unclear how the model handles such directional ambiguity, and this issue is not discussed in the paper.&lt;/li&gt;
        &lt;li&gt;The source of improvement over attention-based architectures is not entirely clear. Since Transformer backbones already capture long-range dependencies through attention, it is unclear what additional benefit the proposed path-based aggregation brings. The paper does not sufficiently clarify the difference between these two mechanisms, nor does it provide analysis to determine whether the gains come from the structural constraint or simply from adding an extra aggregation branch.&lt;/li&gt;
      &lt;/ul&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper presents a well-motivated and technically sound approach to incorporating structural priors into vascular representation learning. The idea of path-based feature aggregation is interesting and offers a fresh perspective, and the method is lightweight, flexible, and shows consistent improvements across datasets and architectures.
That said, the approach relies entirely on implicit supervision, which raises questions about the stability and interpretability of the learned paths. It also does not fully address more complex cases such as bifurcations, and the analysis of the learned structural behavior remains limited.
Overall, the work is solid and shows clear empirical benefits. I lean toward acceptance, though some aspects could be further clarified and strengthened.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;We sincerely thank all reviewers for their positive appreciation:
1.Important problem setting (Meta: “important problem with a practical”; R1: “high practical utility”; R2: “aligns well with … task”; R3: “well-motivated and technically sound”).
2.Lightweight plug-and-play design (Meta: “plug-and-play”; R1: “negligible computational overhead” and “strong architectural flexibility”; R3: “lightweight”).
3.Interesting idea and empirical gains (Meta: “consistent improvements”; R3:“idea is interesting”, “offers a fresh perspective” and “work is solid”).
4.Good clarity (R2:“well-presented and highly clear”).&lt;/p&gt;

  &lt;p&gt;Q1: Clarification of the 3D formulation (Meta, R1)
A1: Due to space limits, we used the 2D formulation to illustrate the shared principle. The 3D version follows the same step-wise path evolution but is not a direct copy: 2D uses turning-angle increments, whereas 3D uses vector directional increments since a scalar angle cannot represent 3D direction changes. Concretely, each voxel predicts an initial direction, bounded step size, and per-step directional updates; directions are updated and normalized, and offsets are accumulated to form the 3D path. Features are then aggregated by 3D grid sampling with trilinear interpolation, with normalized step weighting and residual fusion. For stability, 3D also uses stage-wise step scaling and border padding. We will make this explicit in the revision by adding the 3D formulation, and additional 3D-specific ablation.&lt;/p&gt;

  &lt;p&gt;Q2: Clarification of the Difficulty Gate (Meta, R1, R2)
A2: DG is not trained with explicit difficulty labels; it is learned end-to-end as a utility-driven gate that predicts where structure-guided residuals help segmentation most. Thus, regions where the backbone is weaker tend to receive higher responses, while easier ones are less amplified. A lightweight gate is sufficient because it estimates whether structural guidance is useful from current features, rather than parsing vessels. This is supported by existing evidence: Table2 shows gains over the ungated variant, and Fig.3 shows higher responses around thin branches and bifurcations.&lt;/p&gt;

  &lt;p&gt;Q3: Topology-aware evaluation and p-value (Meta, R1, R2)
A3: To clarify, these results do not involve new experiments; they are post-hoc metrics computed from predictions already generated under the original setting, in response to the reviewers’ request. In the revision, we include Betti-0/Betti-1 evaluation (lower is better) to explicitly assess topology preservation (on DCA1, Betti-0, SegResNet/DSCNet/Ours: 0.788/0.772/0.758; Betti-1: 0.754/0.753/0.730). We also performed a paired Wilcoxon signed-rank test between SegResNet and ours, showing significant differences in Dice(p=0.0045), RDice(p=0.00021), clDice(p=0.00302), Acc(p=0.00696), AUC(p=0.00018), and HD(p=0.00826).&lt;/p&gt;

  &lt;p&gt;Q4: Complementarity to Transformer (Meta, R3)
A4: VSP is complementary to Transformer rather than redundant: attention captures global-based token interactions, while VSP performs explicit structure-aligned propagation along learned anatomical paths, introducing a distinct vascular inductive bias. This is supported by gains on strong backbones such as SwinUNETR (DCA1 Dice:74.33to76.05) with negligible parameter/FLOPs increase, and by consistent improvements across CNN, ViT, and Mamba.&lt;/p&gt;

  &lt;p&gt;Q5: Reliability of VSP path (R3)
A5: VSP path learning is implicit, but risk is controlled by design: the path branch is residual, with bounded step direction updates, normalized weights, and DG gating, limiting harmful perturbations. We agree a single path simplifies bifurcations, but it keeps VSP lightweight and still shows gains there. This choice keeps VSP lightweight; despite this limitation, the visualizations still indicate useful gains around bifurcations. A multi-path extension is a natural future direction, which we will clarify in the revision.&lt;/p&gt;

  &lt;p&gt;We will release our 2D/3D code, and model settings on GitHub, together with the baseline configurations used in our comparisons.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Invite for Rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper addresses an important problem with a practical, plug-and-play module that shows consistent improvements across multiple backbones and datasets. Two reviewers assign Weak Accept (4), one assigns Weak Reject (3). The main concerns are: (1) the path generation mechanism is only defined for 2D, yet applied to 3D datasets without proper formulation or ablation; (2) the Difficulty Gate lacks quantitative validation; (3) topological metrics (Betti) are missing; and (4) the advantage over Transformer backbones is not clarified. These concerns are addressable during rebuttal, though the 3D formulation issue is critical. The paper should proceed to rebuttal with a clear expectation that the authors address these points, particularly the 3D path generation.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper presents a lightweight, plug‑and‑play module (VSP‑Branch) that consistently improves vascular segmentation across CNN, ViT, and Mamba backbones with negligible overhead. Reviewers #1 and #3 gave Weak Accept (4), and Reviewer #2’s Weak Reject (3) concerns (difficulty gate validation, topological metrics, 3D formulation) were adequately addressed in the rebuttal with additional Betti evaluation, statistical testing, and clarified 3D path details. The overall positive consensus and the module’s practical utility support acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-2&quot;&gt;Meta-review #2&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Reject&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper proposes a Vascular Structural Prior Branch (VSP-Branch) that operates at the feature level to generate structure-aware residual updates along learned paths, modulated by a Difficulty Gate, and integrates as a plug-and-play module into CNN, Transformer, and Mamba backbones for vascular segmentation across 2D and 3D datasets. 
After rebuttal, R1 recommends acceptance and R3 remains at weak accept, while R2 maintains reject on the grounds that the Difficulty Gate mechanism is theoretically unjustified and that its qualitative visualization in Figure 3 does not convincingly show focus on genuinely challenging regions. I share R2’s reservations: the network architecture is underspecified, with key components such as the “lightweight convolutional predictor” left without sufficient detail to evaluate the claimed mechanism. I would further add that the formulation does not address how background pixels are processed by the VSP-branch, and how adverse effects from background pixels are mitigated. Given sufficient methodological detail, the work would be impressive, but in its current form, I recommend rejection.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-3&quot;&gt;Meta-review #3&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;After considering the rebuttal and all reviewers’ comments, the overall assessment remains positive. The rebuttal provides reasonable clarification, and the reported experimental results support acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Qi, Yaolei AND Peng, Wenbo AND Wang, Tong AND Xu, Chenwei AND Zhang, Yuan AND Yang, Guanyu</name></author><category term="Body -&gt; Vasculature" /><category term="Modalities -&gt; CT / X-ray" /><category term="Applications -&gt; Image Segmentation" /><category term="Machine Learning -&gt; Deep Learning" /><category term="Qi, Yaolei" /><category term="Peng, Wenbo" /><category term="Wang, Tong" /><category term="Xu, Chenwei" /><category term="Zhang, Yuan" /><category term="Yang, Guanyu" /><summary type="html">Abstract Accurate representation of vascular structures is fundamental for a wide range of clinical applications and remains challenging due to the slender geometry, complex branching topology, and long-range continuity inherent to tree-like vasculature. Although structural priors have been widely explored in vascular image analysis, existing approaches typically integrate such priors implicitly or isotropically during feature aggregation, limiting their ability to faithfully capture directional and structural dependencies along the vessels. In this study, we propose a flexible, structure-guided feature aggregation paradigm that explicitly incorporates anatomical structural cues into the feature aggregation process for efficient vascular representation. The core idea is to aggregate features along directionally guided paths that follow local vascular structures, enabling consistent propagation of structural information while maintaining computational efficiency. To avoid indiscriminate prior enforcement, the proposed paradigm further introduces an adaptive mechanism that selectively injects structural guidance where representation is most challenging. This aggregation paradigm is lightweight, modular, and can be seamlessly integrated into multi-stage feature hierarchies without altering the backbone architecture. Extensive experiments on vascular imaging tasks show that our approach consistently improves the quality of vascular representation. Our code is available at: https://github.com/YaoleiQi/VSP. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/4457_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository https://github.com/YaoleiQi/VSP Link to the Dataset(s) N/A BibTex @InProceedings{QiYao_AFlexible_MICCAI2026,         author = { Qi, Yaolei AND Peng, Wenbo AND Wang, Tong AND Xu, Chenwei AND Zhang, Yuan AND Yang, Guanyu},         title = { { A Flexible Structure-Guided Feature Aggregation Paradigm for Efficient Vascular Representation } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16896},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper The manuscript introduces VSP-Branch, a structure-guided feature aggregation module designed to enhance vascular representation by sampling features along learned anatomical paths and modulating their injection via a Difficulty Gate. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The VSP-Branch demonstrates strong architectural flexibility, functioning as a plug-and-play module that integrates effectively into CNN, ViT, and Mamba-based backbones. 2.The method achieves consistent empirical improvements across a diverse range of imaging modalities and scales, encompassing both 2D (DCA1, DeepRetina) and 3D (COSTA, KIPA) datasets. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.The path construction mechanism lacks formulation and ablation for 3D topologies. The methodology section explicitly defines offset recursion using a standard 2D rotation matrix. However, the framework is applied to 3D datasets (COSTA, KIPA) without detailing how the degrees of freedom (e. g. , pitch and yaw) are parameterized or constrained in 3D space. Furthermore, the ablation study (Table 2) validates the path sampling steps and step-size ranges solely on the 2D DCA1 dataset, leaving the stability of 3D path generation unverified. 2.The functional claim regarding the Difficulty Gate is insufficiently supported by quantitative evidence. The manuscript claims the gate adaptively modulates injection strength in difficult regions based on qualitative visual overlays. Because the gate is supervised only implicitly via the final segmentation objective, there is no quantitative metric to prove it learns to target topological ambiguity rather than defaulting to standard edge detection. 3.The experimental results lack statistical significance testing. While the VSP-Branch consistently improves metrics across baselines, the absolute gains in metrics such as Hausdorff Distance (HD) are frequently marginal (e. g. , HD improvements of less than 0.1 on DCA1). Without p-values, it is difficult to ascertain whether these improvements reflect genuine structural enhancement or mere statistical variance. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The authors claimed to release the source code and/or dataset upon acceptance of the submission. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? This manuscript demonstrates substantial engineering value and a comprehensive empirical workload, despite notable gaps in theoretical formulation. The primary factor driving this score is the high practical utility of the VSP-Branch; it functions as a highly flexible module that achieves consistent performance improvements across diverse baseline architectures, including CNNs, ViTs, and Mamba models, with negligible computational overhead. Furthermore, the authors present a solid validation workload spanning multiple 2D and 3D datasets. However, I withheld a higher score due to specific analytical deficiencies. Most critically, while the method is applied to 3D data, the path generation mechanism is mathematically formulated and ablated strictly within a 2D context using 2D rotation matrices, leaving the 3D topological implementation ambiguous. Additionally, the claim that the Difficulty Gate selectively targets complex anatomical regions relies heavily on qualitative visual overlays rather than rigorous quantitative metrics or statistical significance testing, making it difficult to verify the robustness of marginal metric gains. Overall, the module’s architectural flexibility and strong empirical baseline constitute a net positive contribution to the community that outweighs these presentational shortcomings. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. N/A Review #2 Please describe the contribution of the paper This paper proposes a plug-and-play module called VSP-Branch to explicitly incorporate anatomical cues into the aggregation process, improving vascular representation efficiency. Extensive experiments show VSP-Branch outperforms other methods on Dice metric. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The figures provided in the manuscript are well-presented and highly clear. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.The “Difficulty Gate” is composed of merely a simple convolution. It is unclear how such a straightforward design achieves the authors’ claim of “emphasizing continuity cues more strongly in regions that are difficult to represent while preserving the backbone feature in simpler regions.” Since the Difficulty Gate lacks explicit additional supervision during the training process, how does the module actually learn to distinguish between “simple” and “difficult” regions? 2.Following previous works such as DSCNet, the authors should include additional metrics to evaluate the topological connectivity of the segmented vessels, such as $\beta_0$ and $\beta_1$. 3.In Fig. 3, bounding boxes should be provided on the original full-size images to indicate the cropped areas for the renal dataset visualizations. Furthermore, the visualization of the Difficulty Gate Map is quite confusing. What do the different colored lines represent? The vessels within the areas enclosed by the blue lines appear very clear and do not seem challenging to segment. The authors should provide the baseline segmentation results (“without VSP-Branch”) for comparison. This would help verify whether the VSP-Branch accurately targets and localizes difficult regions in areas that are otherwise prone to mis-segmentation. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The conceptual design of the structure-guided feature aggregation in the VSP-Branch aligns well with the unique characteristics of the vessel segmentation task. However, regarding the core module—the Difficulty Gate Map—both the methodological design and the visualization results in Fig. 3 fail to convincingly demonstrate its claimed role in reinforcing vessel features specifically in difficult regions. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Reject [Post rebuttal] Please justify your final decision from above. While the authors provided quantitative improvements, the core concerns remain unresolved: 1.Unconvincing Mechanism: It remains theoretically unjustified how a simple, unsupervised convolution layer can inherently distinguish between simple and difficult regions. Quantitative gains alone cannot validate a module with questionable underlying logic. 2.Contradictory Visualization: Figure 3 fails to show the Difficulty Gate focusing on genuinely challenging areas (e.g., thin, low-contrast vessels). Furthermore, the authors failed to clarify the meaning of the colored lines, leaving the visualization uninterpretable. Review #3 Please describe the contribution of the paper The paper introduces a structure-guided feature aggregation module, termed VSP-Branch, for improving vascular representation. Rather than relying on conventional local convolutional aggregation or incorporating structural priors only through loss functions, the method aims to explicitly model feature interactions along learned, structure-aligned trajectories. Concretely, it predicts per-pixel directions, step sizes, and turning angles to construct sampling paths, along which features are aggregated to capture long-range structural dependencies. A difficulty-aware gating mechanism is further employed to adaptively control the extent of structural information injected at different spatial locations. The module is lightweight and designed to be easily integrated into various backbone architectures, including CNNs, Transformers, and Mamba-based models. Experimental results across several 2D and 3D vascular datasets show consistent improvements over strong baseline methods. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The paper introduces a representation-level structural aggregation mechanism, which differs from most existing approaches that incorporate structural priors either via specialized convolution operators or loss functions. Modeling feature propagation along learned paths is an interesting and non-trivial design. 2.The proposed VSP-Branch is modular and can be easily plugged into different backbone architectures with negligible computational overhead. This significantly improves its practical applicability. 3.The method is evaluated across multiple datasets and backbone families (CNN, ViT, Mamba), showing consistent performance gains. 4.The Difficulty Gate prevents over-enforcement of structural priors and allows the model to focus on challenging regions such as thin vessels and bifurcations. 5.The paper provides a reasonable interpretation that vascular features exhibit semantic consistency along trajectories, and the proposed method aligns well with this intuition. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. The path construction is learned solely from the segmentation objective without explicit structural supervision. This raises concerns about the reliability of the learned paths, especially in background or vessel-dense regions, where paths may cross background and unintentionally connect nearby vessels. Such behavior could lead to undesired feature mixing and weaken structural consistency, but this issue is not analyzed in the paper. The method predicts a single path for each pixel, which may be limiting in bifurcation regions where multiple valid directions exist. Since vascular structures are inherently tree-like, it is unclear how the model handles such directional ambiguity, and this issue is not discussed in the paper. The source of improvement over attention-based architectures is not entirely clear. Since Transformer backbones already capture long-range dependencies through attention, it is unclear what additional benefit the proposed path-based aggregation brings. The paper does not sufficiently clarify the difference between these two mechanisms, nor does it provide analysis to determine whether the gains come from the structural constraint or simply from adding an extra aggregation branch. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The paper presents a well-motivated and technically sound approach to incorporating structural priors into vascular representation learning. The idea of path-based feature aggregation is interesting and offers a fresh perspective, and the method is lightweight, flexible, and shows consistent improvements across datasets and architectures. That said, the approach relies entirely on implicit supervision, which raises questions about the stability and interpretability of the learned paths. It also does not fully address more complex cases such as bifurcations, and the analysis of the learned structural behavior remains limited. Overall, the work is solid and shows clear empirical benefits. I lean toward acceptance, though some aspects could be further clarified and strengthened. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Author Feedback We sincerely thank all reviewers for their positive appreciation: 1.Important problem setting (Meta: “important problem with a practical”; R1: “high practical utility”; R2: “aligns well with … task”; R3: “well-motivated and technically sound”). 2.Lightweight plug-and-play design (Meta: “plug-and-play”; R1: “negligible computational overhead” and “strong architectural flexibility”; R3: “lightweight”). 3.Interesting idea and empirical gains (Meta: “consistent improvements”; R3:“idea is interesting”, “offers a fresh perspective” and “work is solid”). 4.Good clarity (R2:“well-presented and highly clear”). Q1: Clarification of the 3D formulation (Meta, R1) A1: Due to space limits, we used the 2D formulation to illustrate the shared principle. The 3D version follows the same step-wise path evolution but is not a direct copy: 2D uses turning-angle increments, whereas 3D uses vector directional increments since a scalar angle cannot represent 3D direction changes. Concretely, each voxel predicts an initial direction, bounded step size, and per-step directional updates; directions are updated and normalized, and offsets are accumulated to form the 3D path. Features are then aggregated by 3D grid sampling with trilinear interpolation, with normalized step weighting and residual fusion. For stability, 3D also uses stage-wise step scaling and border padding. We will make this explicit in the revision by adding the 3D formulation, and additional 3D-specific ablation. Q2: Clarification of the Difficulty Gate (Meta, R1, R2) A2: DG is not trained with explicit difficulty labels; it is learned end-to-end as a utility-driven gate that predicts where structure-guided residuals help segmentation most. Thus, regions where the backbone is weaker tend to receive higher responses, while easier ones are less amplified. A lightweight gate is sufficient because it estimates whether structural guidance is useful from current features, rather than parsing vessels. This is supported by existing evidence: Table2 shows gains over the ungated variant, and Fig.3 shows higher responses around thin branches and bifurcations. Q3: Topology-aware evaluation and p-value (Meta, R1, R2) A3: To clarify, these results do not involve new experiments; they are post-hoc metrics computed from predictions already generated under the original setting, in response to the reviewers’ request. In the revision, we include Betti-0/Betti-1 evaluation (lower is better) to explicitly assess topology preservation (on DCA1, Betti-0, SegResNet/DSCNet/Ours: 0.788/0.772/0.758; Betti-1: 0.754/0.753/0.730). We also performed a paired Wilcoxon signed-rank test between SegResNet and ours, showing significant differences in Dice(p=0.0045), RDice(p=0.00021), clDice(p=0.00302), Acc(p=0.00696), AUC(p=0.00018), and HD(p=0.00826). Q4: Complementarity to Transformer (Meta, R3) A4: VSP is complementary to Transformer rather than redundant: attention captures global-based token interactions, while VSP performs explicit structure-aligned propagation along learned anatomical paths, introducing a distinct vascular inductive bias. This is supported by gains on strong backbones such as SwinUNETR (DCA1 Dice:74.33to76.05) with negligible parameter/FLOPs increase, and by consistent improvements across CNN, ViT, and Mamba. Q5: Reliability of VSP path (R3) A5: VSP path learning is implicit, but risk is controlled by design: the path branch is residual, with bounded step direction updates, normalized weights, and DG gating, limiting harmful perturbations. We agree a single path simplifies bifurcations, but it keeps VSP lightweight and still shows gains there. This choice keeps VSP lightweight; despite this limitation, the visualizations still indicate useful gains around bifurcations. A multi-path extension is a natural future direction, which we will clarify in the revision. We will release our 2D/3D code, and model settings on GitHub, together with the baseline configurations used in our comparisons. Meta-Review Meta-review #1 Your recommendation Invite for Rebuttal Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. The paper addresses an important problem with a practical, plug-and-play module that shows consistent improvements across multiple backbones and datasets. Two reviewers assign Weak Accept (4), one assigns Weak Reject (3). The main concerns are: (1) the path generation mechanism is only defined for 2D, yet applied to 3D datasets without proper formulation or ablation; (2) the Difficulty Gate lacks quantitative validation; (3) topological metrics (Betti) are missing; and (4) the advantage over Transformer backbones is not clarified. These concerns are addressable during rebuttal, though the 3D formulation issue is critical. The paper should proceed to rebuttal with a clear expectation that the authors address these points, particularly the 3D path generation. After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. The paper presents a lightweight, plug‑and‑play module (VSP‑Branch) that consistently improves vascular segmentation across CNN, ViT, and Mamba backbones with negligible overhead. Reviewers #1 and #3 gave Weak Accept (4), and Reviewer #2’s Weak Reject (3) concerns (difficulty gate validation, topological metrics, 3D formulation) were adequately addressed in the rebuttal with additional Betti evaluation, statistical testing, and clarified 3D path details. The overall positive consensus and the module’s practical utility support acceptance. Meta-review #2 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Reject Please justify your recommendation. The paper proposes a Vascular Structural Prior Branch (VSP-Branch) that operates at the feature level to generate structure-aware residual updates along learned paths, modulated by a Difficulty Gate, and integrates as a plug-and-play module into CNN, Transformer, and Mamba backbones for vascular segmentation across 2D and 3D datasets. After rebuttal, R1 recommends acceptance and R3 remains at weak accept, while R2 maintains reject on the grounds that the Difficulty Gate mechanism is theoretically unjustified and that its qualitative visualization in Figure 3 does not convincingly show focus on genuinely challenging regions. I share R2’s reservations: the network architecture is underspecified, with key components such as the “lightweight convolutional predictor” left without sufficient detail to evaluate the claimed mechanism. I would further add that the formulation does not address how background pixels are processed by the VSP-branch, and how adverse effects from background pixels are mitigated. Given sufficient methodological detail, the work would be impressive, but in its current form, I recommend rejection. Meta-review #3 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. After considering the rebuttal and all reviewers’ comments, the overall assessment remains positive. The rebuttal provides reasonable clarification, and the reported experimental results support acceptance. back to top</summary></entry><entry><title type="html">A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation</title><link href="https://papers.miccai.org/miccai-2026/0009-Paper5049" rel="alternate" type="text/html" title="A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0009-Paper5049</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0009-Paper5049">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;Delineating the clinical target volume (CTV) in radiotherapy involves complex margins constrained by tumor location and anatomical barriers. While deep learning models automate this process, their rigid reliance on expert-annotated data requires costly retraining whenever clinical guidelines update. To overcome this limitation, we introduce OncoAgent, a novel guideline-aware AI agent framework that seamlessly converts textual clinical guidelines into three-dimensional target contours without any target volume annotation. Evaluated on esophageal cancer cases, the agent achieves a Dice similarity coefficient of 0.842 for the CTV and 0.880 for the planning target volume, demonstrating performance highly comparable to a fully supervised nnU-Net baseline. Notably, in a blinded clinical evaluation, physicians strongly preferred OncoAgent over the supervised baseline, rating it higher in guideline compliance, modification effort, and clinical acceptability. Furthermore, without any retraining, the framework generalizes to alternative esophageal guidelines and shows preliminary extensibility to other anatomical sites (e.g., prostate). Beyond mere volumetric overlap, our agent-based paradigm offers near-instantaneous adaptability to alternative guidelines, providing a scalable and transparent pathway toward interpretability in radiotherapy treatment planning.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/5049_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/5049_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;&lt;a href=&quot;https://github.com/Oncosoft-Research/OncoAgent&quot;&gt;https://github.com/Oncosoft-Research/OncoAgent&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{KimYoo_AGuidelineAware_MICCAI2026,
        author = { Kim, Yoon Jo AND Cho, Wonyoung AND Lee, Jongmin AND Chae, Han Joo AND Park, Hyunki AND Seo, Sang Hoon AND Noh, Jae Myung AND Yang, Kyungmi AND Oh, Dongryul AND Kim, Jin Sung},
        title = { { A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16879},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The main contribution of this paper is an LLM-based agentic framework for Clinical Target Volume (CTV) auto-delineation, validated using esophageal cancer cases. The framework translates free-text clinical guidelines into a structured delineation plan, subsequently, into 3D target volumes. To achieve this, it orchestrates the execution of pre-trained models for Organs at Risk (OARs) segmentation alongside geometric operation tools. 
Using this approach, only free-text clinical guidelines and specific set of parameters, such as body region or dose level, are required as input. Crucially, if clinical guidelines change, the solution does not require costly retraining (unlike traditional deep-learning segmentation methods) but only an update to the guideline descriptions within the prompt. The paper demonstrates the framework adaptability to other protocols and anatomical sites (prostate) in a zero-shot manner. On the esophageal dataset, the framework achieved results comparable to trained nnUNet-based baseline model.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Novel Application: The paper introduces novel application of Large Language Model (LLM) and an agentic framework for Clinical Target Volume (CTV) delineation. A key advantage of this approach is that it bypasses the traditional labor-intensive cycle of data re-annotation and model retraining typically required whenever clinical guidelines are changed. 
2.Clinically Aligned Evaluation: The evaluation is strengthened by the inclusion of expert human assessment, which provides deeper insight than sole standard automated metrics. The findings reveal that while the framework achieves quantitative results comparable to state-of-the art deep learning method, its outputs are much better aligned with human perception. 
3.Safety mechanisms: The framework incorporates a dedicated safety check to ensure the structural validity of generated plans, achieved by implementation of self-refinement mechanism if violations of plan are detected. 
4.Scientific transparency: The authors provide a candid discussion of the study’s limitations. By explicitly addressing factors such as the small evaluation dataset and risk of model hallucinations they show realistic roadmap for future improvements.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;ol&gt;
        &lt;li&gt;Insufficient Implementation Detail: While the proposed concept is compelling and clinically relevant, the paper lacks several critical implementation details. Specifically, the authors do not specify essential hyperparameters such as the temperature settings, or reasoning effort for the LLM. Furthermore, the specific architectures and versions of the models utilized for OARs segmentations are not clearly identified, making it difficult to assess the technical baseline. 
2.Lack of Stochastic and Reliability Analysis: Although the authors acknowledge the risk of model hallucinations the paper provides no quantitative data regarding their frequency. Additionally, there is a lack of information concerning output stability (consistency) when the model is queried multiple times for the same case. Statistics on how often the self-refinement mechanism was triggered are also missing. Such data is important for establishing clinical trust. 
3.Barriers to Reproducibility: The absence of a public code repository or the disclosure of the detailed prompt templates significantly hiders the reproducibility of the study. Given that LLM-based agentic frameworks are sensitive to specific prompting strategies those details are important to be able to validate and build upon the findings.&lt;/li&gt;
      &lt;/ol&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;My recommendation of a Weak Reject is primarily driven by the lack of granular implementation details and the limited analysis regarding system reliability. While the core concept of an LLM-based agentic framework for CTV delineation is both highly relevant and innovative, the current manuscript does not provide sufficient technical depth (e.g., specific hyperparameters and OAR model configurations) or quantitative data on model stability and hallucination rates.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors’ response has successfully addressed my primary concerns, convincing me to upgrade my recommendation from a Weak Reject to an Accept. Specifically, they provided the missing LLM hyperparameters, clarified issues regarding model hallucinations and repeatability, and committed to publicly sharing their prompts and code upon acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents OncoAgent, a guideline-aware agentic framework that converts radiotherapy guidelines into executable tool-call sequences and generates CTV/PTV using pre-trained OAR segmentation models and geometric operations. The problem setting is clinically meaningful, especially given that guideline updates can quickly make purely supervised delineation models outdated. The proposed framework is interesting and potentially useful from an interpretability and adaptability perspective. In the reported experiments, the method achieves performance close to a strong supervised baseline on a small esophageal cancer test set, and receives better blinded physician ratings.&lt;/p&gt;

      &lt;p&gt;That said, I feel several aspects of the work would benefit from clearer positioning and stronger validation. In its current form, the method appears closer to a guideline-to-execution pipeline that translates textual instructions into a predefined sequence of OAR segmentation, GTV expansion, Boolean exclusion, and post-processing, rather than a new end-to-end target delineation model. Some of the stronger claims, such as “training-free,” “zero-shot auto-delineation,” and “cross-site generalization,” may therefore be somewhat overstated relative to the current evidence.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The paper addresses a clinically relevant problem. Frequent guideline changes are indeed challenging for conventional supervised delineation pipelines.
2.The proposed framework has an appealing level of interpretability, since the intermediate execution plan is human-readable and, in principle, reviewable by clinicians.
3.Within the reported experimental setup, the quantitative and qualitative results are promising. OncoAgent performs comparably to nnU-Net(GTV Prior) on CTV/PTV metrics and achieves better physician ratings.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The task formulation would benefit from more precise positioning.
The execution engine operates on the patient CT and existing GTV contours, while relying on pre-trained OAR segmentation models. As such, the method appears closer to guideline-aware target construction given GTV than fully automatic zero-shot target delineation directly from CT. Clarifying this distinction would make the paper easier to interpret.
2.The added value of the LLM/agent component is not yet fully isolated.
At present, the framework seems to mainly parameterize textual guidelines into a largely fixed execution template, rather than perform genuinely complex clinical reasoning. The reported average of 1.13 LLM inference calls per case also suggests that the process may be highly templated. Without a non-LLM baseline, such as a manually implemented rule engine or template-based parser using the same geometric pipeline, it is difficult to determine how much of the gain comes from the agentic component itself versus the explicit rule-based construction.
3.It is not yet fully demonstrated that the framework can accommodate the full complexity of esophageal target expansion rules.
In practice, esophageal CTV delineation often involves conditional, hierarchical, and context-dependent decisions, and may not always be reducible to simple margin expansion followed by OAR subtraction. The manuscript notes that guideline ranges are resolved using patient-specific context, such as dose level or physician preferences, but this decision mechanism is not described in enough detail to assess its robustness and reproducibility.
4.The validation remains limited in scale.
The study includes 40 patients in total, with only 8 test cases, and the blinded physician evaluation is based on those same 8 cases rated by 2 radiation oncologists. For a paper emphasizing clinical preference and clinical competitiveness, this is still a relatively limited validation.
5.The evidence for cross-guideline and cross-site generalization is still indirect.
The alternative-guideline and prostate experiments are mainly supported by Tool Call F1, which reflects plan-sequence agreement rather than actual contour quality or downstream clinical utility. In particular, the prostate result of 0.64 suggests that substantial challenges remain when extending the framework to a new anatomical site.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Satisfactory&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper addresses a clinically relevant problem, and the motivation is clear. I think the idea of making target delineation more adaptable to changing guidelines is interesting. The reported results are promising in the presented setting. However, I remain slightly below the acceptance threshold for three main reasons: first, the method would benefit from clearer positioning, since it relies on existing GTV contours and pre-trained OAR segmentation and therefore feels closer to guideline-aware target construction than fully automatic zero-shot delineation; second, the added value of the LLM/agent component is not yet fully isolated from the underlying rule-based geometric pipeline; and third, the current evidence is still limited in scale and in the strength of the cross-guideline/cross-site validation. Overall, I think this is a promising direction, but the current version still needs clearer framing and stronger validation.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Reject&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;My main concern after rebuttal is the insufficient specification of the patient-specific decision mechanism. The authors state that the LLM can resolve guideline ranges and context-dependent rules using patient-specific information, but it remains unclear how GTV location is extracted, how upper/mid/lower esophageal involvement is determined, how margin ranges and physician preferences are operationalized, and whether such decisions are correct at the case level. These are not minor implementation details, because they directly affect the generated CTV/PTV and are central to the claimed guideline-aware reasoning capability.&lt;/p&gt;

      &lt;p&gt;The rebuttal provides illustrative examples, such as selecting nodal regions according to tumor location, but it does not provide a sufficiently reproducible mechanism for how these decisions are made from the available inputs. It also remains unclear whether the human-readable tool-call plan is only an audit trail after the LLM decision has been made, or whether there is a systematic validation process to ensure that the upstream clinical reasoning is correct before execution. In this sense, interpretability of the execution plan does not fully address the reliability of the patient-specific decision process.&lt;/p&gt;

      &lt;p&gt;I appreciate the authors’ clarification that “zero-shot” means no CTV-specific training and their willingness to tone down overstrong claims. However, the current evidence still does not fully demonstrate that the framework can robustly handle conditional, hierarchical, and context-dependent clinical rules in a reproducible way. Since these decisions directly determine the final contours, this remains a central limitation. I therefore maintain my reject assessment.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper proposes OncoAgent, a guideline-aware AI agent that converts textual radiotherapy guidelines into 3D target volume contours in a zero-shot manner. The framework uses a large language model to translate guideline text into structured tool-call sequences, which are executed using pre-trained OAR segmentation models and geometric operations (e.g., dilation and subtraction). Evaluated on esophageal cancer cases, OncoAgent achieves performance comparable to a supervised nnU-Net baseline while demonstrating improved physician preference in a blinded evaluation. The approach further claims zero-shot adaptability to alternative guidelines and anatomical sites.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The work introduces a fundamentally different approach to target delineation by explicitly leveraging clinical guidelines rather than relying solely on data-driven learning. This addresses a real limitation of current deep learning models, particularly their inability to adapt to evolving clinical protocols without retraining.
2.The formulation (e.g., CTV = dilation of GTV minus OARs) closely reflects actual clinical reasoning, making the method intuitive and potentially more trustworthy.
3.The method achieves performance comparable to a strong supervised baseline (nnU-Net with GTV prior), despite not requiring task-specific training for CTV delineation.
4.The clinical assessment is a significant strength and provides valuable insight beyond standard segmentation metrics.
5.The explicit, human-readable planning steps improve transparency and enable rapid adaptation to new guidelines without retraining.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.While the framework does not require CTV-specific training, it relies on pre-trained OAR segmentation models and a large language model. This should be more clearly articulated to avoid overstating the training-free nature of the approach.
2.The method critically depends on pre-trained OAR segmentation models, yet the specific models used, their training data, and their performance are not described. Since OAR delineation directly influences the final CTV, this dependency should be explicitly characterized.
3.The evaluation is performed on a small test set (n=8), and the clinical assessment involves only two physicians. This limits the robustness of both quantitative and subjective conclusions.
4.The Likert-based evaluation aggregates ratings across a small number of cases and ratersr. Additionally, the absence of ground-truth contours in the evaluation makes it difficult to contextualize physician preferences.
5.While the method is applied to alternative guidelines and anatomical sites, quantitative evaluation beyond the primary esophageal task is limited. For example, prostate results are reported only via tool-call metrics rather than volumetric accuracy.
6.As acknowledged by the authors, the framework may be susceptible to misinterpretation or hallucination by the LLM, which could lead to incorrect delineation steps.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper presents an innovative and clinically relevant direction for guideline-aware radiotherapy planning. To further strengthen the work, the authors could improve clarity around model dependencies, expand clinical validation, and provide more detailed evaluation of generalization across anatomical sites. Including ground-truth contours in the physician evaluation would also help contextualize subjective ratings.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper introduces a novel and clinically meaningful paradigm for target volume delineation that directly incorporates clinical guidelines into the contouring process. The approach demonstrates competitive performance with supervised methods while offering improved interpretability and adaptability. Despite limitations in dataset size, evaluation detail, and clarity of certain claims, the overall contribution is significant and has the potential to influence future research directions in radiotherapy planning.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The rebuttal satisfactorily addresses several of my concerns. In particular, the authors clarify that zero-shot refers to the absence of CTV-specific training rather than complete independence from pretrained models, provide additional details on the OAR segmentation model, and appropriately acknowledge the limited prostate generalization claim. The concerns regarding small validation scale and limited physician evaluation remain, but these are reasonable limitations given the novelty and clinical relevance of the proposed paradigm. I therefore maintain my positive recommendation.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;We thank the reviewers for their thoughtful and constructive feedback. We are grateful for the recognition of OncoAgent’s novelty as the first guideline-aware AI agent for CTV delineation and its clinical relevance. We address the main concerns below.&lt;/p&gt;

  &lt;p&gt;1.Implementation Details and Reproducibility (R1, R3)
We will add the following sentence after the implementation paragraph in Sec. 3.1: “GPT-5.2 was configured with temperature=1.0 and reasoning_effort=none; a modified 3D U-Net trained on a total of 2,870 CT and MR volumes with Dice 0.80–0.98 was used as an OAR segmentation model.” The full system prompt, JSON schemas, and execution code will be released on GitHub upon acceptance.&lt;/p&gt;

  &lt;p&gt;2.Analysis on Hallucination Frequency (R1, R3)
OncoAgent’s hallucination risk falls into two categories: invalid tool-call plans (structural) and misinterpreted guidelines (semantic). For structural hallucination, the reported 1.13 calls/case in Sec. 3.1 implies that ~87% of plans pass schema validation on first generation while ~13% trigger self-refinement; all executed plans are structurally valid by construction. For semantic correctness, Tool Call F1 of 0.73–1.00 across heterogeneous esophageal guidelines (Table 3) provides indirect evidence. Regarding output stability, we empirically observed that the schema-validation and self-refinement loop drove repeated queries on the same case toward consistent, schema-valid plans. We will clarify these points in Sec. 3.1.
3.Task Formulation (R2, R3)
By “zero-shot” we mean no CTV-specific training is needed; the framework still relies on a pre-trained LLM and OAR segmentation models. To avoid overstatement, we will rephrase terms such as “training-free” as “without any CTV annotation”.&lt;/p&gt;

  &lt;p&gt;4.Value of LLM Compared to Rule Engine and Complexity of CTV Decision Rules (R2)
The contribution of the LLM is not measured by the number of LLM API calls, but by the elimination of per-guideline engineering cost. Target volume plans differ across guidelines—e.g., IJROBP follows a GTV→CTV→PTV pipeline with anatomy-aware CTV expansions, whereas CROSS expands GTV directly to PTV with fixed geometric margins. The value of LLM lies in autonomously adapting to different guidelines without any system-prompt or code modification, whereas rule engines would require hand-crafted parsers per guideline.
Regarding CTV decision complexity, Equation 1 represents the computational skeleton shared across all guidelines, while guideline-specific complexity and patient-specific context are handled through LLM reasoning. For example, the LLM can dynamically determine which lymph node regions to include based on GTV location—cervical nodes for upper esophageal lesions, celiac nodes for lower. Furthermore, for cases requiring careful adjustment such as re-irradiation, our human-readable plan supports human-in-the-loop audit before contour generation.&lt;/p&gt;

  &lt;p&gt;5.Cross-Generalization (R2, R3)
We acknowledge that the prostate Tool Call F1 of 0.64 indicates limited generalization to this site, and we will tone down the corresponding cross-site claims in Sec. 3.4.While this work focuses on esophageal cancer, the framework is designed to be extensible: incorporating site-specific features (e.g., OAR contours, target-construction methods) would enable broader applicability. We also clarify that Tool Call F1 is not a loose proxy: since the ground-truth calls were constructed by a human and execution is deterministic, matching sequence implies matching CTV contours.&lt;/p&gt;

  &lt;p&gt;6.Validation Scale (R2, R3)
The current validation scale is acknowledged in our existing limitations, with larger cross-institutional, cross-guideline validation as a future direction. We note that, although ground truth per physician was unavailable, the expert contours in Fig. 3 served as a common visual reference during evaluation. Within this scope, OncoAgent demonstrates a shift from learning-from-data to reasoning-from-guidelines, evidencing a scalable and auditable pathway.&lt;/p&gt;
&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Invite for Rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper receives 2 weak reject and 1 weak accept. All reviewers raise several valid major concerns. Rebuttal is invited.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Reject&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper still receives mixed reviews after rebuttal. One reviewer increases the score from negative to positive. Although this direction is interesting, I agree with R2 that, under the agent system, it is still not clear how GTV location is extracted, how upper/mid/lower esophageal involvement is determined, how margin ranges and physician preferences are operationalized, and whether such decisions are correct at the case level. The current quantitative performance is inferior to GTV prior based nnUNet. Moreover, the comparing method is not state-of-the-art, as there is more sophisticated clinical target volume segmentation utilizing GTV, OARs in esophageal and head &amp;amp; neck cancers. The evaluation dataset size is also small. Hence, I lean to rejection of this work.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-2&quot;&gt;Meta-review #2&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The LLM driven agent framework for delineation while considering the organs at risk is novel. Authors carefully addressed most of reviewers’ concerns in the rebuttal. However, some of the limitations including lack of cross-site validation, limited data size, and how model is refined for individual patient-specific decision making must be clarified in the final paper.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-3&quot;&gt;Meta-review #3&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents an LLM-based agentic framework that translates free-text clinical guidelines into an executable target volume contouring plan. I agree with the reviewers’ consensus regarding the clinical relevance of this work. After the rebuttal phase, the authors generally addressed the primary technical concerns raised during the review process. While some concerns remain (such as the issues regarding the small dataset scale, system reliability, and reproducibility), I think the novelty and clinical relevance of the proposed paradigm still outweigh the limitations, and it could provide valuable insights and inspiration for future research in automated radiotherapy planning. Therefore, I recommend accept.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Kim, Yoon Jo AND Cho, Wonyoung AND Lee, Jongmin AND Chae, Han Joo AND Park, Hyunki AND Seo, Sang Hoon AND Noh, Jae Myung AND Yang, Kyungmi AND Oh, Dongryul AND Kim, Jin Sung</name></author><category term="Body -&gt; Lung / Thoracic" /><category term="Modalities -&gt; CT / X-ray" /><category term="Applications -&gt; Image Segmentation" /><category term="Machine Learning -&gt; Multimodal Models / LLMs / VLMs" /><category term="Kim, Yoon Jo" /><category term="Cho, Wonyoung" /><category term="Lee, Jongmin" /><category term="Chae, Han Joo" /><category term="Park, Hyunki" /><category term="Seo, Sang Hoon" /><category term="Noh, Jae Myung" /><category term="Yang, Kyungmi" /><category term="Oh, Dongryul" /><category term="Kim, Jin Sung" /><summary type="html">Abstract Delineating the clinical target volume (CTV) in radiotherapy involves complex margins constrained by tumor location and anatomical barriers. While deep learning models automate this process, their rigid reliance on expert-annotated data requires costly retraining whenever clinical guidelines update. To overcome this limitation, we introduce OncoAgent, a novel guideline-aware AI agent framework that seamlessly converts textual clinical guidelines into three-dimensional target contours without any target volume annotation. Evaluated on esophageal cancer cases, the agent achieves a Dice similarity coefficient of 0.842 for the CTV and 0.880 for the planning target volume, demonstrating performance highly comparable to a fully supervised nnU-Net baseline. Notably, in a blinded clinical evaluation, physicians strongly preferred OncoAgent over the supervised baseline, rating it higher in guideline compliance, modification effort, and clinical acceptability. Furthermore, without any retraining, the framework generalizes to alternative esophageal guidelines and shows preliminary extensibility to other anatomical sites (e.g., prostate). Beyond mere volumetric overlap, our agent-based paradigm offers near-instantaneous adaptability to alternative guidelines, providing a scalable and transparent pathway toward interpretability in radiotherapy treatment planning. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/5049_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository https://github.com/Oncosoft-Research/OncoAgent Link to the Dataset(s) N/A BibTex @InProceedings{KimYoo_AGuidelineAware_MICCAI2026,         author = { Kim, Yoon Jo AND Cho, Wonyoung AND Lee, Jongmin AND Chae, Han Joo AND Park, Hyunki AND Seo, Sang Hoon AND Noh, Jae Myung AND Yang, Kyungmi AND Oh, Dongryul AND Kim, Jin Sung},         title = { { A Guideline-Aware AI Agent for Zero-Shot Target Volume Auto-Delineation } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16879},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper The main contribution of this paper is an LLM-based agentic framework for Clinical Target Volume (CTV) auto-delineation, validated using esophageal cancer cases. The framework translates free-text clinical guidelines into a structured delineation plan, subsequently, into 3D target volumes. To achieve this, it orchestrates the execution of pre-trained models for Organs at Risk (OARs) segmentation alongside geometric operation tools. Using this approach, only free-text clinical guidelines and specific set of parameters, such as body region or dose level, are required as input. Crucially, if clinical guidelines change, the solution does not require costly retraining (unlike traditional deep-learning segmentation methods) but only an update to the guideline descriptions within the prompt. The paper demonstrates the framework adaptability to other protocols and anatomical sites (prostate) in a zero-shot manner. On the esophageal dataset, the framework achieved results comparable to trained nnUNet-based baseline model. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.Novel Application: The paper introduces novel application of Large Language Model (LLM) and an agentic framework for Clinical Target Volume (CTV) delineation. A key advantage of this approach is that it bypasses the traditional labor-intensive cycle of data re-annotation and model retraining typically required whenever clinical guidelines are changed. 2.Clinically Aligned Evaluation: The evaluation is strengthened by the inclusion of expert human assessment, which provides deeper insight than sole standard automated metrics. The findings reveal that while the framework achieves quantitative results comparable to state-of-the art deep learning method, its outputs are much better aligned with human perception. 3.Safety mechanisms: The framework incorporates a dedicated safety check to ensure the structural validity of generated plans, achieved by implementation of self-refinement mechanism if violations of plan are detected. 4.Scientific transparency: The authors provide a candid discussion of the study’s limitations. By explicitly addressing factors such as the small evaluation dataset and risk of model hallucinations they show realistic roadmap for future improvements. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. Insufficient Implementation Detail: While the proposed concept is compelling and clinically relevant, the paper lacks several critical implementation details. Specifically, the authors do not specify essential hyperparameters such as the temperature settings, or reasoning effort for the LLM. Furthermore, the specific architectures and versions of the models utilized for OARs segmentations are not clearly identified, making it difficult to assess the technical baseline. 2.Lack of Stochastic and Reliability Analysis: Although the authors acknowledge the risk of model hallucinations the paper provides no quantitative data regarding their frequency. Additionally, there is a lack of information concerning output stability (consistency) when the model is queried multiple times for the same case. Statistics on how often the self-refinement mechanism was triggered are also missing. Such data is important for establishing clinical trust. 3.Barriers to Reproducibility: The absence of a public code repository or the disclosure of the detailed prompt templates significantly hiders the reproducibility of the study. Given that LLM-based agentic frameworks are sensitive to specific prompting strategies those details are important to be able to validate and build upon the findings. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? My recommendation of a Weak Reject is primarily driven by the lack of granular implementation details and the limited analysis regarding system reliability. While the core concept of an LLM-based agentic framework for CTV delineation is both highly relevant and innovative, the current manuscript does not provide sufficient technical depth (e.g., specific hyperparameters and OAR model configurations) or quantitative data on model stability and hallucination rates. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The authors’ response has successfully addressed my primary concerns, convincing me to upgrade my recommendation from a Weak Reject to an Accept. Specifically, they provided the missing LLM hyperparameters, clarified issues regarding model hallucinations and repeatability, and committed to publicly sharing their prompts and code upon acceptance. Review #2 Please describe the contribution of the paper This paper presents OncoAgent, a guideline-aware agentic framework that converts radiotherapy guidelines into executable tool-call sequences and generates CTV/PTV using pre-trained OAR segmentation models and geometric operations. The problem setting is clinically meaningful, especially given that guideline updates can quickly make purely supervised delineation models outdated. The proposed framework is interesting and potentially useful from an interpretability and adaptability perspective. In the reported experiments, the method achieves performance close to a strong supervised baseline on a small esophageal cancer test set, and receives better blinded physician ratings. That said, I feel several aspects of the work would benefit from clearer positioning and stronger validation. In its current form, the method appears closer to a guideline-to-execution pipeline that translates textual instructions into a predefined sequence of OAR segmentation, GTV expansion, Boolean exclusion, and post-processing, rather than a new end-to-end target delineation model. Some of the stronger claims, such as “training-free,” “zero-shot auto-delineation,” and “cross-site generalization,” may therefore be somewhat overstated relative to the current evidence. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The paper addresses a clinically relevant problem. Frequent guideline changes are indeed challenging for conventional supervised delineation pipelines. 2.The proposed framework has an appealing level of interpretability, since the intermediate execution plan is human-readable and, in principle, reviewable by clinicians. 3.Within the reported experimental setup, the quantitative and qualitative results are promising. OncoAgent performs comparably to nnU-Net(GTV Prior) on CTV/PTV metrics and achieves better physician ratings. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.The task formulation would benefit from more precise positioning. The execution engine operates on the patient CT and existing GTV contours, while relying on pre-trained OAR segmentation models. As such, the method appears closer to guideline-aware target construction given GTV than fully automatic zero-shot target delineation directly from CT. Clarifying this distinction would make the paper easier to interpret. 2.The added value of the LLM/agent component is not yet fully isolated. At present, the framework seems to mainly parameterize textual guidelines into a largely fixed execution template, rather than perform genuinely complex clinical reasoning. The reported average of 1.13 LLM inference calls per case also suggests that the process may be highly templated. Without a non-LLM baseline, such as a manually implemented rule engine or template-based parser using the same geometric pipeline, it is difficult to determine how much of the gain comes from the agentic component itself versus the explicit rule-based construction. 3.It is not yet fully demonstrated that the framework can accommodate the full complexity of esophageal target expansion rules. In practice, esophageal CTV delineation often involves conditional, hierarchical, and context-dependent decisions, and may not always be reducible to simple margin expansion followed by OAR subtraction. The manuscript notes that guideline ranges are resolved using patient-specific context, such as dose level or physician preferences, but this decision mechanism is not described in enough detail to assess its robustness and reproducibility. 4.The validation remains limited in scale. The study includes 40 patients in total, with only 8 test cases, and the blinded physician evaluation is based on those same 8 cases rated by 2 radiation oncologists. For a paper emphasizing clinical preference and clinical competitiveness, this is still a relatively limited validation. 5.The evidence for cross-guideline and cross-site generalization is still indirect. The alternative-guideline and prostate experiments are mainly supported by Tool Call F1, which reflects plan-sequence agreement rather than actual contour quality or downstream clinical utility. In particular, the prostate result of 0.64 suggests that substantial challenges remain when extending the framework to a new anatomical site. Please rate the clarity and organization of this paper Satisfactory Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The paper addresses a clinically relevant problem, and the motivation is clear. I think the idea of making target delineation more adaptable to changing guidelines is interesting. The reported results are promising in the presented setting. However, I remain slightly below the acceptance threshold for three main reasons: first, the method would benefit from clearer positioning, since it relies on existing GTV contours and pre-trained OAR segmentation and therefore feels closer to guideline-aware target construction than fully automatic zero-shot delineation; second, the added value of the LLM/agent component is not yet fully isolated from the underlying rule-based geometric pipeline; and third, the current evidence is still limited in scale and in the strength of the cross-guideline/cross-site validation. Overall, I think this is a promising direction, but the current version still needs clearer framing and stronger validation. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Reject [Post rebuttal] Please justify your final decision from above. My main concern after rebuttal is the insufficient specification of the patient-specific decision mechanism. The authors state that the LLM can resolve guideline ranges and context-dependent rules using patient-specific information, but it remains unclear how GTV location is extracted, how upper/mid/lower esophageal involvement is determined, how margin ranges and physician preferences are operationalized, and whether such decisions are correct at the case level. These are not minor implementation details, because they directly affect the generated CTV/PTV and are central to the claimed guideline-aware reasoning capability. The rebuttal provides illustrative examples, such as selecting nodal regions according to tumor location, but it does not provide a sufficiently reproducible mechanism for how these decisions are made from the available inputs. It also remains unclear whether the human-readable tool-call plan is only an audit trail after the LLM decision has been made, or whether there is a systematic validation process to ensure that the upstream clinical reasoning is correct before execution. In this sense, interpretability of the execution plan does not fully address the reliability of the patient-specific decision process. I appreciate the authors’ clarification that “zero-shot” means no CTV-specific training and their willingness to tone down overstrong claims. However, the current evidence still does not fully demonstrate that the framework can robustly handle conditional, hierarchical, and context-dependent clinical rules in a reproducible way. Since these decisions directly determine the final contours, this remains a central limitation. I therefore maintain my reject assessment. Review #3 Please describe the contribution of the paper The paper proposes OncoAgent, a guideline-aware AI agent that converts textual radiotherapy guidelines into 3D target volume contours in a zero-shot manner. The framework uses a large language model to translate guideline text into structured tool-call sequences, which are executed using pre-trained OAR segmentation models and geometric operations (e.g., dilation and subtraction). Evaluated on esophageal cancer cases, OncoAgent achieves performance comparable to a supervised nnU-Net baseline while demonstrating improved physician preference in a blinded evaluation. The approach further claims zero-shot adaptability to alternative guidelines and anatomical sites. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The work introduces a fundamentally different approach to target delineation by explicitly leveraging clinical guidelines rather than relying solely on data-driven learning. This addresses a real limitation of current deep learning models, particularly their inability to adapt to evolving clinical protocols without retraining. 2.The formulation (e.g., CTV = dilation of GTV minus OARs) closely reflects actual clinical reasoning, making the method intuitive and potentially more trustworthy. 3.The method achieves performance comparable to a strong supervised baseline (nnU-Net with GTV prior), despite not requiring task-specific training for CTV delineation. 4.The clinical assessment is a significant strength and provides valuable insight beyond standard segmentation metrics. 5.The explicit, human-readable planning steps improve transparency and enable rapid adaptation to new guidelines without retraining. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.While the framework does not require CTV-specific training, it relies on pre-trained OAR segmentation models and a large language model. This should be more clearly articulated to avoid overstating the training-free nature of the approach. 2.The method critically depends on pre-trained OAR segmentation models, yet the specific models used, their training data, and their performance are not described. Since OAR delineation directly influences the final CTV, this dependency should be explicitly characterized. 3.The evaluation is performed on a small test set (n=8), and the clinical assessment involves only two physicians. This limits the robustness of both quantitative and subjective conclusions. 4.The Likert-based evaluation aggregates ratings across a small number of cases and ratersr. Additionally, the absence of ground-truth contours in the evaluation makes it difficult to contextualize physician preferences. 5.While the method is applied to alternative guidelines and anatomical sites, quantitative evaluation beyond the primary esophageal task is limited. For example, prostate results are reported only via tool-call metrics rather than volumetric accuracy. 6.As acknowledged by the authors, the framework may be susceptible to misinterpretation or hallucination by the LLM, which could lead to incorrect delineation steps. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html The paper presents an innovative and clinically relevant direction for guideline-aware radiotherapy planning. To further strengthen the work, the authors could improve clarity around model dependencies, expand clinical validation, and provide more detailed evaluation of generalization across anatomical sites. Including ground-truth contours in the physician evaluation would also help contextualize subjective ratings. Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The paper introduces a novel and clinically meaningful paradigm for target volume delineation that directly incorporates clinical guidelines into the contouring process. The approach demonstrates competitive performance with supervised methods while offering improved interpretability and adaptability. Despite limitations in dataset size, evaluation detail, and clarity of certain claims, the overall contribution is significant and has the potential to influence future research directions in radiotherapy planning. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The rebuttal satisfactorily addresses several of my concerns. In particular, the authors clarify that zero-shot refers to the absence of CTV-specific training rather than complete independence from pretrained models, provide additional details on the OAR segmentation model, and appropriately acknowledge the limited prostate generalization claim. The concerns regarding small validation scale and limited physician evaluation remain, but these are reasonable limitations given the novelty and clinical relevance of the proposed paradigm. I therefore maintain my positive recommendation. Author Feedback We thank the reviewers for their thoughtful and constructive feedback. We are grateful for the recognition of OncoAgent’s novelty as the first guideline-aware AI agent for CTV delineation and its clinical relevance. We address the main concerns below. 1.Implementation Details and Reproducibility (R1, R3) We will add the following sentence after the implementation paragraph in Sec. 3.1: “GPT-5.2 was configured with temperature=1.0 and reasoning_effort=none; a modified 3D U-Net trained on a total of 2,870 CT and MR volumes with Dice 0.80–0.98 was used as an OAR segmentation model.” The full system prompt, JSON schemas, and execution code will be released on GitHub upon acceptance. 2.Analysis on Hallucination Frequency (R1, R3) OncoAgent’s hallucination risk falls into two categories: invalid tool-call plans (structural) and misinterpreted guidelines (semantic). For structural hallucination, the reported 1.13 calls/case in Sec. 3.1 implies that ~87% of plans pass schema validation on first generation while ~13% trigger self-refinement; all executed plans are structurally valid by construction. For semantic correctness, Tool Call F1 of 0.73–1.00 across heterogeneous esophageal guidelines (Table 3) provides indirect evidence. Regarding output stability, we empirically observed that the schema-validation and self-refinement loop drove repeated queries on the same case toward consistent, schema-valid plans. We will clarify these points in Sec. 3.1. 3.Task Formulation (R2, R3) By “zero-shot” we mean no CTV-specific training is needed; the framework still relies on a pre-trained LLM and OAR segmentation models. To avoid overstatement, we will rephrase terms such as “training-free” as “without any CTV annotation”. 4.Value of LLM Compared to Rule Engine and Complexity of CTV Decision Rules (R2) The contribution of the LLM is not measured by the number of LLM API calls, but by the elimination of per-guideline engineering cost. Target volume plans differ across guidelines—e.g., IJROBP follows a GTV→CTV→PTV pipeline with anatomy-aware CTV expansions, whereas CROSS expands GTV directly to PTV with fixed geometric margins. The value of LLM lies in autonomously adapting to different guidelines without any system-prompt or code modification, whereas rule engines would require hand-crafted parsers per guideline. Regarding CTV decision complexity, Equation 1 represents the computational skeleton shared across all guidelines, while guideline-specific complexity and patient-specific context are handled through LLM reasoning. For example, the LLM can dynamically determine which lymph node regions to include based on GTV location—cervical nodes for upper esophageal lesions, celiac nodes for lower. Furthermore, for cases requiring careful adjustment such as re-irradiation, our human-readable plan supports human-in-the-loop audit before contour generation. 5.Cross-Generalization (R2, R3) We acknowledge that the prostate Tool Call F1 of 0.64 indicates limited generalization to this site, and we will tone down the corresponding cross-site claims in Sec. 3.4.While this work focuses on esophageal cancer, the framework is designed to be extensible: incorporating site-specific features (e.g., OAR contours, target-construction methods) would enable broader applicability. We also clarify that Tool Call F1 is not a loose proxy: since the ground-truth calls were constructed by a human and execution is deterministic, matching sequence implies matching CTV contours. 6.Validation Scale (R2, R3) The current validation scale is acknowledged in our existing limitations, with larger cross-institutional, cross-guideline validation as a future direction. We note that, although ground truth per physician was unavailable, the expert contours in Fig. 3 served as a common visual reference during evaluation. Within this scope, OncoAgent demonstrates a shift from learning-from-data to reasoning-from-guidelines, evidencing a scalable and auditable pathway. Meta-Review Meta-review #1 Your recommendation Invite for Rebuttal Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. The paper receives 2 weak reject and 1 weak accept. All reviewers raise several valid major concerns. Rebuttal is invited. After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Reject Please justify your recommendation. This paper still receives mixed reviews after rebuttal. One reviewer increases the score from negative to positive. Although this direction is interesting, I agree with R2 that, under the agent system, it is still not clear how GTV location is extracted, how upper/mid/lower esophageal involvement is determined, how margin ranges and physician preferences are operationalized, and whether such decisions are correct at the case level. The current quantitative performance is inferior to GTV prior based nnUNet. Moreover, the comparing method is not state-of-the-art, as there is more sophisticated clinical target volume segmentation utilizing GTV, OARs in esophageal and head &amp;amp; neck cancers. The evaluation dataset size is also small. Hence, I lean to rejection of this work. Meta-review #2 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. The LLM driven agent framework for delineation while considering the organs at risk is novel. Authors carefully addressed most of reviewers’ concerns in the rebuttal. However, some of the limitations including lack of cross-site validation, limited data size, and how model is refined for individual patient-specific decision making must be clarified in the final paper. Meta-review #3 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. This paper presents an LLM-based agentic framework that translates free-text clinical guidelines into an executable target volume contouring plan. I agree with the reviewers’ consensus regarding the clinical relevance of this work. After the rebuttal phase, the authors generally addressed the primary technical concerns raised during the review process. While some concerns remain (such as the issues regarding the small dataset scale, system reliability, and reproducibility), I think the novelty and clinical relevance of the proposed paradigm still outweigh the limitations, and it could provide valuable insights and inspiration for future research in automated radiotherapy planning. Therefore, I recommend accept. back to top</summary></entry><entry><title type="html">A Heterogeneous Prognosis Prediction Framework with Global Brain Connectivity and Local Image Features</title><link href="https://papers.miccai.org/miccai-2026/0010-Paper1042" rel="alternate" type="text/html" title="A Heterogeneous Prognosis Prediction Framework with Global Brain Connectivity and Local Image Features" /><published>2026-09-21T00:00:00-04:00</published><updated>2026-09-21T00:00:00-04:00</updated><id>https://papers.miccai.org/miccai-2026/0010-Paper1042</id><content type="html" xml:base="https://papers.miccai.org/miccai-2026/0010-Paper1042">&lt;h1 id=&quot;abstract-id&quot;&gt;Abstract&lt;/h1&gt;
&lt;p&gt;Diffuse glioma is the most prevalent malignant brain tumor, and accurate prognosis prediction is essential in personalized treatment to improve outcomes. Although numerous deep learning based methods have been proposed for prognosis prediction, most of them rely solely on local image features of tumor regions. They neglect the fact that the brain is an integrated and highly interconnected system, and tumors can disrupt the global brain connectivity, which is closely associated with prognosis. To address this limitation, we present a novel heterogeneous framework that integrates global brain connectivity and local image features for prognosis prediction. Specifically, our framework consists of a DTI based graph neural network (GNN) and a structural MR based local feature network (LFN), which are coupled through an iterative pipeline. In each iteration, the GNN models the global connectivity among brain regions (nodes) with consideration of brain compensation caused by tumor disruptions, while the LFN leverages the resulting global brain connectivity to extract effective local image features from each brain region, which are used to refine the brain nodes in the GNN for improved connectivity modeling in the next iteration. Through the iterative process, the global brain connectivity in the GNN and the local image features in the LFN are mutually enhanced, leading to accurate prognosis prediction. Extensive experiments on a public UCSF dataset of 493 diffuse glioma patients demonstrate that our framework outperforms the state-of-the-art methods. Further ablation studies reveal that the global brain connectivity is crucial and highly related to prognosis.
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;link-id&quot;&gt;Links to Paper and Supplementary Materials&lt;/h1&gt;
&lt;p&gt;Main Paper (Open Access Version): &lt;a href=&quot;https://papers.miccai.org/miccai-2026/paper/1042_paper.pdf&quot; target=&quot;_blank&quot;&gt;https://papers.miccai.org/miccai-2026/paper/1042_paper.pdf&lt;/a&gt;&lt;/p&gt;

&lt;p&gt;SharedIt Link: Not yet available&lt;/p&gt;

&lt;p&gt;SpringerLink (DOI): Not yet available&lt;/p&gt;

&lt;p&gt;Supplementary Material: Not Submitted
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;code-id&quot;&gt;Link to the Code Repository&lt;/h1&gt;
&lt;p&gt;N/A
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;dataset-id&quot;&gt;Link to the Dataset(s)&lt;/h1&gt;
&lt;p&gt;Cam-CAN dataset: &lt;a href=&quot;https://cam-can.mrc-cbu.cam.ac.uk/dataset/&quot;&gt;https://cam-can.mrc-cbu.cam.ac.uk/dataset/&lt;/a&gt;
UCSF-PDGM dataset: &lt;a href=&quot;https://www.cancerimagingarchive.net/collection/ucsf-pdgm/&quot;&gt;https://www.cancerimagingarchive.net/collection/ucsf-pdgm/&lt;/a&gt;
&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;h1 id=&quot;bibtex-id&quot;&gt;BibTex&lt;/h1&gt;
&lt;pre&gt;&lt;code class=&quot;language-{verbatim}&quot;&gt;@InProceedings{LinJin_AHeterogeneous_MICCAI2026,
        author = { Lin, Jingfeng AND Pan, Junjun AND Wang, Jinda AND Tang, Zhenyu},
        title = { { A Heterogeneous Prognosis Prediction Framework with Global Brain Connectivity and Local Image Features } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16887},
        month = {September},
        page = {pending}
}

&lt;/code&gt;&lt;/pre&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;review-id&quot;&gt;Reviews&lt;/h1&gt;

&lt;h3 id=&quot;review-1&quot;&gt;Review #1&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents a novel heterogeneous framework for predicting the prognosis of diffuse glioma by integrating global brain connectivity derived from DTI with local image features extracted from structural MRI. The methodology couples a GNN with a LFN in an iterative pipeline, explicitly modeling the brain’s compensatory mechanisms in response to tumor disruption. The proposed framework is evaluated on the public UCSF dataset and demonstrates superior performance compared to 14 state-of-the-art methods across multiple metrics. The work addresses a valid clinical limitation—the neglect of global brain network disruption in current deep learning models—and provides a well-engineered solution to fuse complementary modalities. While the technical execution is solid and the results are promising, there are concerns regarding methodological clarity, generalizability, and computational justification that prevent a higher score.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The explicit modeling of the brain’s compensatory mechanism via a modified adjacency matrix and sparsity loss is a compelling insight. It bridges neuro-oncology principles with graph learning design, moving beyond standard feature concatenation. 
The mutual refinement between the GNN (global connectivity) and LFN (local regional features) is a strong architectural contribution. The visualization of the attention matrix convergence across iterations supports the claim that the two streams enhance each other. 
The evaluation against 14 SOTA methods, including both radiomics，deep learning and GNN-based approaches, provides a convincing benchmark. The ablation studies effectively isolate the contribution of the compensation mechanism and the iterative pipeline.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The framework introduces several specific design choices that lack empirical or theoretical justification. For instance, the threshold for defining tumor-invaded regions is arbitrarily set at a voxel ratio of &amp;gt;40%, and the number of iterations is fixed at six based on “empirical observations. “ The paper does not discuss the sensitivity of the model to these choices nor the computational overhead of unrolling the model six times during training and inference. 
Regarding the CamCAN healthy reference: You use it to compute &amp;amp;A_{healthy} . Have you tested the model’s performance when using the average connectivity of the non-tumor hemispheres of the UCSF patients themselves as a pseudo-healthy reference? 
In the revision, please elaborate on Equation (7). A diagram or detailed description of how the 90x90 attention matrix transforms the 3D feature volumes of 90 regions would significantly improve the clarity of the method section.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper offers a conceptually strong contribution by coupling global brain connectivity with local imaging features through an iterative GNN-LFN framework and explicitly modeling compensatory reorganization, which is clinically novel. The comprehensive benchmarking against 14 methods on a public cohort is commendable. However, the work is held back by several unresolved issues: arbitrary parameter choices (e.g., fixed six iterations and a 40% tumor threshold) lack sensitivity analysis, the heavy reliance on an external healthy control dataset (CamCAN) raises practical generalizability concerns, and the technical description of connectivity-guided feature fusion (Equation 7) remains opaque. A convincing rebuttal that clarifies these points would solidify acceptance; otherwise, I would not object to rejection.&lt;/p&gt;

    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The main reason is that the paper presents a well-motivated multimodal prognostic framework with a clear degree of novelty and promising empirical performance. The integration of structural connectomics with local MRI features is clinically meaningful, and the iterative mutual refinement scheme is more interesting than standard multimodal fusion. The benchmarking is extensive, and the rebuttal successfully addresses several concerns about external reference data, structural priors, fairness of baselines, and computational practicality.
Although some concerns are not fully eliminated, they have been sufficiently reduced such that they no longer outweigh the paper’s strengths. In particular, the authors now provide a more credible justification for the use of Cam-CAN, explain why the contralateral hemisphere is not an appropriate substitute, and support the compensatory prior with an alternative-prior comparison. These clarifications improve my confidence that the proposed design choices are thoughtful rather than arbitrary.
I still believe the revised manuscript should more clearly present the sensitivity analyses and improve the explanation of the connectivity-guided feature fusion mechanism. However, these now seem like fixable presentation and completeness issues, rather than flaws that undermine the central contribution. On balance, I think the paper is marginally above the acceptance threshold.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-2&quot;&gt;Review #2&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;This paper presents a heterogeneous framework for diffuse glioma prognosis prediction that couples DTI-driven structural connectivity with regional features extracted from structural MRI in an iterative pipeline, and further introduces a compensation-aware graph design for Cox-based survival risk prediction.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The paper presents a relatively novel architecture that fuses structural connectivity and region-level imaging features through an iterative coupling scheme, where global connectivity guides local representations and refined local features are fed back to update graph nodes.
The paper also introduces clinically and biologically motivated priors into prognosis modeling by explicitly considering whole-brain network disruption and a compensation-related mechanism in diffuse glioma.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Unfair and incomplete baseline comparison.
The experimental comparison is not sufficiently fair to support the claimed contribution. The paper mainly contrasts local image feature (LIF) methods and global brain connectivity (GBC) methods, but does not include multimodal baselines that combine DTI-based structural networks with structural MRI features under comparable settings. In particular, the manuscript lacks simple but strong baselines such as direct early/late fusion of a DTI graph branch and an sMRI encoder branch, cross-attention fusion, or a unified heterogeneous graph model without the compensation prior. Without such comparisons, it is difficult to determine whether the performance gain truly comes from the proposed core idea, or simply from using more modalities, more modules, and stronger inductive bias.&lt;/p&gt;

      &lt;p&gt;2.The claimed mechanism is not clearly validated.
The contribution and functional role of the proposed compensation mechanism are not convincingly established. The so-called compensatory mechanism appears more like a strong hand-crafted prior than a mechanism learned from data and rigorously verified. The model explicitly rewrites the graph structure by enforcing connections between tumor-invaded regions and peritumoral/contralateral regions, and further imposes sparsity on tumor-node attention to favor a limited set of compensatory regions. As a result, the mechanism is injected into the model rather than discovered by it. More importantly, the paper does not show that this prior is necessary or correct for prognosis prediction: there are no comparisons with weaker or more generic structural priors, no stability analysis across different tumor thresholds, compensation definitions, or sparsity weights, and no direct evidence that the gain comes specifically from modeling compensation rather than from graph modification in general.&lt;/p&gt;

      &lt;p&gt;3.Key hyperparameters and design choices are insufficiently justified.
Several important methodological choices appear arbitrary and are not adequately motivated. For example, the paper directly adopts AAL-90 for parcellation, uses PANDA to construct white matter connectivity from DTI, binarizes edges with a fixed threshold of 0.2, and defines tumor-invaded regions using a tumor voxel ratio greater than 40%. However, the manuscript does not explain how these thresholds were selected, whether they were inherited from prior work, tuned on training data, or chosen heuristically. No sensitivity analysis is provided, which weakens the credibility and reproducibility of the method.&lt;/p&gt;

      &lt;p&gt;4.The writing and presentation are weak.
The manuscript is difficult to follow in several places. The introduction does not clearly develop the motivation for the design, and the method section does not cleanly separate the high-level architecture from implementation details and hyperparameter choices. Some definitions are also unclear or inconsistent; for example, the specific image encoder used for local feature extraction is not clearly described, and the notation around Eq. (8) is confusing (What is A_ori?). The experimental section is also poorly organized, with dataset description, baselines, and evaluation metrics presented in a mixed manner. In addition, the result analysis is relatively shallow, and Fig. 3 is not easy to interpret.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Poor&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not provide sufficient information for reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(2) Reject — should be rejected, independent of rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;While the paper explores an interesting and clinically motivated direction, I do not think the current version provides sufficiently strong evidence to support its main claims. The iterative coupling of structural connectivity and regional imaging features is a meaningful idea, and the attempt to incorporate whole-brain network disruption and compensation-related priors is a positive aspect.&lt;/p&gt;

      &lt;p&gt;That said, my main concerns are about rigor and validation. The baseline comparison is not fully convincing, since strong modality-matched multimodal baselines are missing; several key design choices and hyperparameters are insufficiently justified; and the proposed compensation mechanism is not yet validated clearly enough to show that the observed gains truly come from this specific modeling idea. In addition, the presentation could be improved in terms of clarity and organization. Overall, I felt the paper has potential, but in its current form it does not yet meet the standard for acceptance.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Very confident (4)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h3 id=&quot;review-3&quot;&gt;Review #3&lt;/h3&gt;

&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Please describe the contribution of the paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors propose a heterogeneous deep learning framework for the prognosis prediction of diffuse glioma. The framework integrates global brain connectivity (derived from DTI) via a Graph Attention Network (GAT) and local image features (derived from multimodal sMR) via a Local Feature Network (LFN). These two networks are coupled iteratively to mutually refine node features and graph topology. The authors also introduce a “compensatory mechanism” by modifying the adjacency matrix to force connections between tumor-invaded regions and peritumoral/contralateral regions.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.The idea of iteratively coupling structural MRI (local features) and DTI (global connectivity) is conceptually elegant and well-motivated for brain tumor analysis.
2.The authors evaluate their method against a robust set of 14 state-of-the-art methods, showing strong quantitative improvements across multiple metrics.&lt;/p&gt;
      &lt;ol&gt;
        &lt;li&gt;The ablation study clearly demonstrates the added value of the iterative pipeline and the compensation module.&lt;/li&gt;
      &lt;/ol&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;1.Diffuse gliomas frequently cause severe structural deformation (mass effect), shifting midline structures and distorting normal anatomy. Standard linear or non-linear registration to a healthy atlas (like AAL) fails drastically in the presence of large tumors.
2.Equation (1) explicitly hardcodes connections between tumor-invaded regions and peritumoral/contralateral regions. The model highlights peritumoral and contralateral regions because the authors manually wired the graph to prioritize them, not necessarily because the network organically discovered a biological compensatory mechanism.
3.Cam-CAN is a completely different dataset with likely different acquisition protocols, scanner types, and preprocessing pipelines compared to the UCSF glioma dataset. The authors do not address how they harmonize the DTI data between UCSF and Cam-CAN.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please rate the clarity and organization of this paper&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Good&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;N/A&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation. What were the major factors that led you to your overall score for this paper?&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The methodological flaw of computing patient-specific connectivity alterations by subtracting a healthy baseline from a completely different dataset/scanner invalidates the core input of the model. The results are highly likely driven by confounding site effects.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Reviewer confidence&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Confident but not absolutely certain (3)&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;[Post rebuttal] Please justify your final decision from above.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;The authors have satisfactorily addressed these concerns:
1.The clarification that ComBat harmonization and age-matching were applied during preprocessing significantly alleviates my concerns regarding confounding site effects. This is a standard and robust approach.
2.The use of ANTs-based nonlinear registration with lesion masking is the correct methodological choice for dealing with mass effect in glioma patients.
3.The authors provided a compelling justification for their structural prior by comparing it to a fully connected strategy, demonstrating a significant performance drop (C-index 0.663 vs 0.717) when the model is left to learn connections without the clinical prior.&lt;/p&gt;

      &lt;p&gt;I am raising my score to a “Weak Accept” on the strict condition that all these clarifications are explicitly integrated into the final camera-ready version of the paper, as the method is not reproducible or scientifically sound without them.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;h2&gt;&lt;br /&gt;&lt;br /&gt;&lt;/h2&gt;
&lt;h1 id=&quot;authorFeedback-id&quot;&gt;Author Feedback&lt;/h1&gt;
&lt;blockquote&gt;
  &lt;p&gt;We thank the reviewers for recognizing the novelty of our framework (R1–R3). We respond to the key concerns regarding methodology and validation as follows:&lt;/p&gt;

  &lt;p&gt;Q1.The reliance on the external Cam-CAN dataset and potential domain shift/harmonization issues:
To mitigate potential domain shift between the Cam-CAN and UCSF datasets, we applied ComBat harmonization during preprocessing, which is a widely used strategy for multi-site data harmonization (Solanes et al., NeuroImage 2023). Furthermore, the selected Cam-CAN subjects were age-matched to the UCSF cohort (55.91 ± 17.71 vs. 56.80 ± 15.06 years), reducing confounding from age-related connectome changes.&lt;/p&gt;

  &lt;p&gt;Q2.Why not use the contralateral non-tumor hemisphere as a pseudo-healthy reference?
Given that glioma can induce whole-brain network alterations, the contralateral hemisphere may also exhibit connectivity disruption or compensatory reorganization, making it unsuitable as a reliable pseudo-healthy reference.&lt;/p&gt;

  &lt;p&gt;Q3.Robustness of atlas registration to tumor mass effect:
To reduce the influence of tumor mass effect on atlas registration, we performed ANTs-based nonlinear registration with lesion masking (Wei et al., TMI 2023). This reduces the contribution of tumor-related abnormal regions to the similarity metric during optimization, making the registration primarily driven by non-lesioned brain tissues.&lt;/p&gt;

  &lt;p&gt;Q4.Lack of fair, modality-matched multimodal baselines:
Our comparison includes modality-matched multimodal baselines: DC-R2SNs, IDH-GNN, and MaskGNN share our modality setting, with node features from sMRI and edge/connectivity features from DTI.&lt;/p&gt;

  &lt;p&gt;Q5.Necessity of the compensatory structural prior:
The compensatory structural prior reduces the model search space, thereby improving performance, and this design is supported by clinical evidence (Duffau, Cortex 2014). We have also tested alternative structural priors, including a more flexible fully connected strategy in which tumor regions were connected to all normal brain regions, allowing the model to automatically learn potential compensatory connections. However, this strategy only achieved a C-index of 0.663, lower than 0.717 obtained with the compensatory structural prior. This may be because the fully connected strategy introduces many biologically less meaningful connections and aggravates graph over-smoothing, thereby reducing performance.&lt;/p&gt;

  &lt;p&gt;Q6.Hyperparameter sensitivity, training/inference time, and reproducibility:
Based on clinical evidence that cross-regional compensation is less likely to occur when tumor involvement is below 30% or above 60% (Duffau, Cortex 2014), we tested tumor-voxel thresholds within the 30%–60% range. Model performance remained overall stable across this range, with 40% slightly outperforming the other thresholds.
For the number of unrolling iterations, we observed that the attention matrix and brain-region features generally converged within 6 steps; further increasing the number of iterations only introduced additional computational cost without improving model performance. 
Under Ubuntu 22.04 with an RTX 3090 GPU, the training and inference times were 26.95 min per epoch and 41.9 s, respectively. 
To ensure reproducibility, we will release the code upon acceptance.&lt;/p&gt;

&lt;/blockquote&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;

&lt;hr /&gt;
&lt;h1 id=&quot;metareview-id&quot;&gt;Meta-Review&lt;/h1&gt;

&lt;h2 id=&quot;meta-review-1&quot;&gt;Meta-review #1&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;Your recommendation&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Invite for Rebuttal&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your decision.  In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Although reviewers appreciated the novel GNN-LFN framework integrating DTI and MRI for glioma prognosis (score: 4), critical concerns regarding methodology and validation warrant a rebuttal (scores: 2, 3). To support acceptance, the authors must address: (1) the reliance on the external Cam-CAN dataset and potential domain shift/harmonization issues; (2) the robustness of atlas registration against tumor mass effects; (3) the lack of fair, modality-matched multimodal baselines; and (4) justifications and sensitivity analyses for key handcrafted priors and hyperparameters (e.g., tumor threshold, iterations).&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;As strongly advised by Reviewers 1 and 3, acceptance is conditional upon explicitly incorporating these critical clarifications, along with the hyperparameter sensitivity analyses discussed in the rebuttal, into the camera-ready manuscript to guarantee scientific rigor and reproducibility.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-2&quot;&gt;Meta-review #2&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;After carefully reviewing the authors’ rebuttal, I have revised my initial assessment to Accept. The authors have provided a thorough, evidence-based response that directly addresses the core concerns raised by all three reviewers.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;&lt;/p&gt;
&lt;h2 id=&quot;meta-review-3&quot;&gt;Meta-review #3&lt;/h2&gt;
&lt;ul&gt;
  &lt;li&gt;&lt;strong&gt;After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;Accept&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
  &lt;li&gt;&lt;strong&gt;Please justify your recommendation.&lt;/strong&gt;
    &lt;blockquote&gt;
      &lt;p&gt;I recommend acceptance. Although the paper is borderline, the rebuttal substantially addressed the major concerns regarding domain shift, lesion-aware registration, modality-matched baselines, and the compensation prior. Two reviewers now support acceptance, and the remaining issues are mainly about presentation and completeness rather than core validity. The camera-ready version should explicitly include the rebuttal clarifications.&lt;/p&gt;
    &lt;/blockquote&gt;
  &lt;/li&gt;
&lt;/ul&gt;

&lt;p&gt;&lt;br /&gt;&lt;br /&gt;
&lt;a href=&quot;&quot;&gt;&lt;b&gt;back to top&lt;/b&gt;&lt;/a&gt;&lt;/p&gt;

&lt;hr /&gt;</content><author><name>Lin, Jingfeng AND Pan, Junjun AND Wang, Jinda AND Tang, Zhenyu</name></author><category term="Body -&gt; Brain" /><category term="Modalities -&gt; CT / X-ray" /><category term="Modalities -&gt; Diffusion MRI" /><category term="Modalities -&gt; Functional MRI" /><category term="Applications -&gt; Outcome Prediction / Prognosis / Longitudinal Modeling" /><category term="Machine Learning -&gt; Deep Learning" /><category term="Lin, Jingfeng" /><category term="Pan, Junjun" /><category term="Wang, Jinda" /><category term="Tang, Zhenyu" /><summary type="html">Abstract Diffuse glioma is the most prevalent malignant brain tumor, and accurate prognosis prediction is essential in personalized treatment to improve outcomes. Although numerous deep learning based methods have been proposed for prognosis prediction, most of them rely solely on local image features of tumor regions. They neglect the fact that the brain is an integrated and highly interconnected system, and tumors can disrupt the global brain connectivity, which is closely associated with prognosis. To address this limitation, we present a novel heterogeneous framework that integrates global brain connectivity and local image features for prognosis prediction. Specifically, our framework consists of a DTI based graph neural network (GNN) and a structural MR based local feature network (LFN), which are coupled through an iterative pipeline. In each iteration, the GNN models the global connectivity among brain regions (nodes) with consideration of brain compensation caused by tumor disruptions, while the LFN leverages the resulting global brain connectivity to extract effective local image features from each brain region, which are used to refine the brain nodes in the GNN for improved connectivity modeling in the next iteration. Through the iterative process, the global brain connectivity in the GNN and the local image features in the LFN are mutually enhanced, leading to accurate prognosis prediction. Extensive experiments on a public UCSF dataset of 493 diffuse glioma patients demonstrate that our framework outperforms the state-of-the-art methods. Further ablation studies reveal that the global brain connectivity is crucial and highly related to prognosis. Links to Paper and Supplementary Materials Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/1042_paper.pdf SharedIt Link: Not yet available SpringerLink (DOI): Not yet available Supplementary Material: Not Submitted Link to the Code Repository N/A Link to the Dataset(s) Cam-CAN dataset: https://cam-can.mrc-cbu.cam.ac.uk/dataset/ UCSF-PDGM dataset: https://www.cancerimagingarchive.net/collection/ucsf-pdgm/ BibTex @InProceedings{LinJin_AHeterogeneous_MICCAI2026,         author = { Lin, Jingfeng AND Pan, Junjun AND Wang, Jinda AND Tang, Zhenyu},         title = { { A Heterogeneous Prognosis Prediction Framework with Global Brain Connectivity and Local Image Features } },         booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},         year = {2026}, publisher = {Springer Nature Switzerland},         volume = {LNCS 16887},         month = {September},        page = {pending} } Reviews Review #1 Please describe the contribution of the paper This paper presents a novel heterogeneous framework for predicting the prognosis of diffuse glioma by integrating global brain connectivity derived from DTI with local image features extracted from structural MRI. The methodology couples a GNN with a LFN in an iterative pipeline, explicitly modeling the brain’s compensatory mechanisms in response to tumor disruption. The proposed framework is evaluated on the public UCSF dataset and demonstrates superior performance compared to 14 state-of-the-art methods across multiple metrics. The work addresses a valid clinical limitation—the neglect of global brain network disruption in current deep learning models—and provides a well-engineered solution to fuse complementary modalities. While the technical execution is solid and the results are promising, there are concerns regarding methodological clarity, generalizability, and computational justification that prevent a higher score. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The explicit modeling of the brain’s compensatory mechanism via a modified adjacency matrix and sparsity loss is a compelling insight. It bridges neuro-oncology principles with graph learning design, moving beyond standard feature concatenation. The mutual refinement between the GNN (global connectivity) and LFN (local regional features) is a strong architectural contribution. The visualization of the attention matrix convergence across iterations supports the claim that the two streams enhance each other. The evaluation against 14 SOTA methods, including both radiomics，deep learning and GNN-based approaches, provides a convincing benchmark. The ablation studies effectively isolate the contribution of the compensation mechanism and the iterative pipeline. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. The framework introduces several specific design choices that lack empirical or theoretical justification. For instance, the threshold for defining tumor-invaded regions is arbitrarily set at a voxel ratio of &amp;gt;40%, and the number of iterations is fixed at six based on “empirical observations. “ The paper does not discuss the sensitivity of the model to these choices nor the computational overhead of unrolling the model six times during training and inference. Regarding the CamCAN healthy reference: You use it to compute &amp;amp;A_{healthy} . Have you tested the model’s performance when using the average connectivity of the non-tumor hemispheres of the UCSF patients themselves as a pseudo-healthy reference? In the revision, please elaborate on Equation (7). A diagram or detailed description of how the 90x90 attention matrix transforms the 3D feature volumes of 90 regions would significantly improve the clarity of the method section. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The paper offers a conceptually strong contribution by coupling global brain connectivity with local imaging features through an iterative GNN-LFN framework and explicitly modeling compensatory reorganization, which is clinically novel. The comprehensive benchmarking against 14 methods on a public cohort is commendable. However, the work is held back by several unresolved issues: arbitrary parameter choices (e.g., fixed six iterations and a 40% tumor threshold) lack sensitivity analysis, the heavy reliance on an external healthy control dataset (CamCAN) raises practical generalizability concerns, and the technical description of connectivity-guided feature fusion (Equation 7) remains opaque. A convincing rebuttal that clarifies these points would solidify acceptance; otherwise, I would not object to rejection. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The main reason is that the paper presents a well-motivated multimodal prognostic framework with a clear degree of novelty and promising empirical performance. The integration of structural connectomics with local MRI features is clinically meaningful, and the iterative mutual refinement scheme is more interesting than standard multimodal fusion. The benchmarking is extensive, and the rebuttal successfully addresses several concerns about external reference data, structural priors, fairness of baselines, and computational practicality. Although some concerns are not fully eliminated, they have been sufficiently reduced such that they no longer outweigh the paper’s strengths. In particular, the authors now provide a more credible justification for the use of Cam-CAN, explain why the contralateral hemisphere is not an appropriate substitute, and support the compensatory prior with an alternative-prior comparison. These clarifications improve my confidence that the proposed design choices are thoughtful rather than arbitrary. I still believe the revised manuscript should more clearly present the sensitivity analyses and improve the explanation of the connectivity-guided feature fusion mechanism. However, these now seem like fixable presentation and completeness issues, rather than flaws that undermine the central contribution. On balance, I think the paper is marginally above the acceptance threshold. Review #2 Please describe the contribution of the paper This paper presents a heterogeneous framework for diffuse glioma prognosis prediction that couples DTI-driven structural connectivity with regional features extracted from structural MRI in an iterative pipeline, and further introduces a compensation-aware graph design for Cox-based survival risk prediction. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. The paper presents a relatively novel architecture that fuses structural connectivity and region-level imaging features through an iterative coupling scheme, where global connectivity guides local representations and refined local features are fed back to update graph nodes. The paper also introduces clinically and biologically motivated priors into prognosis modeling by explicitly considering whole-brain network disruption and a compensation-related mechanism in diffuse glioma. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.Unfair and incomplete baseline comparison. The experimental comparison is not sufficiently fair to support the claimed contribution. The paper mainly contrasts local image feature (LIF) methods and global brain connectivity (GBC) methods, but does not include multimodal baselines that combine DTI-based structural networks with structural MRI features under comparable settings. In particular, the manuscript lacks simple but strong baselines such as direct early/late fusion of a DTI graph branch and an sMRI encoder branch, cross-attention fusion, or a unified heterogeneous graph model without the compensation prior. Without such comparisons, it is difficult to determine whether the performance gain truly comes from the proposed core idea, or simply from using more modalities, more modules, and stronger inductive bias. 2.The claimed mechanism is not clearly validated. The contribution and functional role of the proposed compensation mechanism are not convincingly established. The so-called compensatory mechanism appears more like a strong hand-crafted prior than a mechanism learned from data and rigorously verified. The model explicitly rewrites the graph structure by enforcing connections between tumor-invaded regions and peritumoral/contralateral regions, and further imposes sparsity on tumor-node attention to favor a limited set of compensatory regions. As a result, the mechanism is injected into the model rather than discovered by it. More importantly, the paper does not show that this prior is necessary or correct for prognosis prediction: there are no comparisons with weaker or more generic structural priors, no stability analysis across different tumor thresholds, compensation definitions, or sparsity weights, and no direct evidence that the gain comes specifically from modeling compensation rather than from graph modification in general. 3.Key hyperparameters and design choices are insufficiently justified. Several important methodological choices appear arbitrary and are not adequately motivated. For example, the paper directly adopts AAL-90 for parcellation, uses PANDA to construct white matter connectivity from DTI, binarizes edges with a fixed threshold of 0.2, and defines tumor-invaded regions using a tumor voxel ratio greater than 40%. However, the manuscript does not explain how these thresholds were selected, whether they were inherited from prior work, tuned on training data, or chosen heuristically. No sensitivity analysis is provided, which weakens the credibility and reproducibility of the method. 4.The writing and presentation are weak. The manuscript is difficult to follow in several places. The introduction does not clearly develop the motivation for the design, and the method section does not cleanly separate the high-level architecture from implementation details and hyperparameter choices. Some definitions are also unclear or inconsistent; for example, the specific image encoder used for local feature extraction is not clearly described, and the notation around Eq. (8) is confusing (What is A_ori?). The experimental section is also poorly organized, with dataset description, baselines, and evaluation metrics presented in a mixed manner. In addition, the result analysis is relatively shallow, and Fig. 3 is not easy to interpret. Please rate the clarity and organization of this paper Poor Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not provide sufficient information for reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (2) Reject — should be rejected, independent of rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? While the paper explores an interesting and clinically motivated direction, I do not think the current version provides sufficiently strong evidence to support its main claims. The iterative coupling of structural connectivity and regional imaging features is a meaningful idea, and the attempt to incorporate whole-brain network disruption and compensation-related priors is a positive aspect. That said, my main concerns are about rigor and validation. The baseline comparison is not fully convincing, since strong modality-matched multimodal baselines are missing; several key design choices and hyperparameters are insufficiently justified; and the proposed compensation mechanism is not yet validated clearly enough to show that the observed gains truly come from this specific modeling idea. In addition, the presentation could be improved in terms of clarity and organization. Overall, I felt the paper has potential, but in its current form it does not yet meet the standard for acceptance. Reviewer confidence Very confident (4) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. N/A [Post rebuttal] Please justify your final decision from above. N/A Review #3 Please describe the contribution of the paper The authors propose a heterogeneous deep learning framework for the prognosis prediction of diffuse glioma. The framework integrates global brain connectivity (derived from DTI) via a Graph Attention Network (GAT) and local image features (derived from multimodal sMR) via a Local Feature Network (LFN). These two networks are coupled iteratively to mutually refine node features and graph topology. The authors also introduce a “compensatory mechanism” by modifying the adjacency matrix to force connections between tumor-invaded regions and peritumoral/contralateral regions. Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting. 1.The idea of iteratively coupling structural MRI (local features) and DTI (global connectivity) is conceptually elegant and well-motivated for brain tumor analysis. 2.The authors evaluate their method against a robust set of 14 state-of-the-art methods, showing strong quantitative improvements across multiple metrics. The ablation study clearly demonstrates the added value of the iterative pipeline and the compensation module. Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work. 1.Diffuse gliomas frequently cause severe structural deformation (mass effect), shifting midline structures and distorting normal anatomy. Standard linear or non-linear registration to a healthy atlas (like AAL) fails drastically in the presence of large tumors. 2.Equation (1) explicitly hardcodes connections between tumor-invaded regions and peritumoral/contralateral regions. The model highlights peritumoral and contralateral regions because the authors manually wired the graph to prioritize them, not necessarily because the network organically discovered a biological compensatory mechanism. 3.Cam-CAN is a completely different dataset with likely different acquisition protocols, scanner types, and preprocessing pipelines compared to the UCSF glioma dataset. The authors do not address how they harmonize the DTI data between UCSF and Cam-CAN. Please rate the clarity and organization of this paper Good Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance. The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility. Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation? N/A Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html N/A Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making. (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal Please justify your recommendation. What were the major factors that led you to your overall score for this paper? The methodological flaw of computing patient-specific connectivity alterations by subtracting a healthy baseline from a completely different dataset/scanner invalidates the core input of the model. The results are highly likely driven by confounding site effects. Reviewer confidence Confident but not absolutely certain (3) [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper. Accept [Post rebuttal] Please justify your final decision from above. The authors have satisfactorily addressed these concerns: 1.The clarification that ComBat harmonization and age-matching were applied during preprocessing significantly alleviates my concerns regarding confounding site effects. This is a standard and robust approach. 2.The use of ANTs-based nonlinear registration with lesion masking is the correct methodological choice for dealing with mass effect in glioma patients. 3.The authors provided a compelling justification for their structural prior by comparing it to a fully connected strategy, demonstrating a significant performance drop (C-index 0.663 vs 0.717) when the model is left to learn connections without the clinical prior. I am raising my score to a “Weak Accept” on the strict condition that all these clarifications are explicitly integrated into the final camera-ready version of the paper, as the method is not reproducible or scientifically sound without them. Author Feedback We thank the reviewers for recognizing the novelty of our framework (R1–R3). We respond to the key concerns regarding methodology and validation as follows: Q1.The reliance on the external Cam-CAN dataset and potential domain shift/harmonization issues: To mitigate potential domain shift between the Cam-CAN and UCSF datasets, we applied ComBat harmonization during preprocessing, which is a widely used strategy for multi-site data harmonization (Solanes et al., NeuroImage 2023). Furthermore, the selected Cam-CAN subjects were age-matched to the UCSF cohort (55.91 ± 17.71 vs. 56.80 ± 15.06 years), reducing confounding from age-related connectome changes. Q2.Why not use the contralateral non-tumor hemisphere as a pseudo-healthy reference? Given that glioma can induce whole-brain network alterations, the contralateral hemisphere may also exhibit connectivity disruption or compensatory reorganization, making it unsuitable as a reliable pseudo-healthy reference. Q3.Robustness of atlas registration to tumor mass effect: To reduce the influence of tumor mass effect on atlas registration, we performed ANTs-based nonlinear registration with lesion masking (Wei et al., TMI 2023). This reduces the contribution of tumor-related abnormal regions to the similarity metric during optimization, making the registration primarily driven by non-lesioned brain tissues. Q4.Lack of fair, modality-matched multimodal baselines: Our comparison includes modality-matched multimodal baselines: DC-R2SNs, IDH-GNN, and MaskGNN share our modality setting, with node features from sMRI and edge/connectivity features from DTI. Q5.Necessity of the compensatory structural prior: The compensatory structural prior reduces the model search space, thereby improving performance, and this design is supported by clinical evidence (Duffau, Cortex 2014). We have also tested alternative structural priors, including a more flexible fully connected strategy in which tumor regions were connected to all normal brain regions, allowing the model to automatically learn potential compensatory connections. However, this strategy only achieved a C-index of 0.663, lower than 0.717 obtained with the compensatory structural prior. This may be because the fully connected strategy introduces many biologically less meaningful connections and aggravates graph over-smoothing, thereby reducing performance. Q6.Hyperparameter sensitivity, training/inference time, and reproducibility: Based on clinical evidence that cross-regional compensation is less likely to occur when tumor involvement is below 30% or above 60% (Duffau, Cortex 2014), we tested tumor-voxel thresholds within the 30%–60% range. Model performance remained overall stable across this range, with 40% slightly outperforming the other thresholds. For the number of unrolling iterations, we observed that the attention matrix and brain-region features generally converged within 6 steps; further increasing the number of iterations only introduced additional computational cost without improving model performance. Under Ubuntu 22.04 with an RTX 3090 GPU, the training and inference times were 26.95 min per epoch and 41.9 s, respectively. To ensure reproducibility, we will release the code upon acceptance. Meta-Review Meta-review #1 Your recommendation Invite for Rebuttal Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal. Although reviewers appreciated the novel GNN-LFN framework integrating DTI and MRI for glioma prognosis (score: 4), critical concerns regarding methodology and validation warrant a rebuttal (scores: 2, 3). To support acceptance, the authors must address: (1) the reliance on the external Cam-CAN dataset and potential domain shift/harmonization issues; (2) the robustness of atlas registration against tumor mass effects; (3) the lack of fair, modality-matched multimodal baselines; and (4) justifications and sensitivity analyses for key handcrafted priors and hyperparameters (e.g., tumor threshold, iterations). After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. As strongly advised by Reviewers 1 and 3, acceptance is conditional upon explicitly incorporating these critical clarifications, along with the hyperparameter sensitivity analyses discussed in the rebuttal, into the camera-ready manuscript to guarantee scientific rigor and reproducibility. Meta-review #2 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. After carefully reviewing the authors’ rebuttal, I have revised my initial assessment to Accept. The authors have provided a thorough, evidence-based response that directly addresses the core concerns raised by all three reviewers. Meta-review #3 After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal. Accept Please justify your recommendation. I recommend acceptance. Although the paper is borderline, the rebuttal substantially addressed the major concerns regarding domain shift, lesion-aware registration, modality-matched baselines, and the compensation prior. Two reviewers now support acceptance, and the remaining issues are mainly about presentation and completeness rather than core validity. The camera-ready version should explicitly include the rebuttal clarifications. back to top</summary></entry></feed>