Abstract

Cross-modality medical image synthesis remains fundamentally challenging due to residual spatial inconsistencies from complex non-rigid organ deformations and localized anatomical variability that persist beyond registration. Existing approaches either rely on explicit pre-registration procedures or integrate registration-guided training modules that are not retained during inference, thereby limiting their ability to model residual misalignments and anatomy-specific deformation dynamics. We propose a unified framework that combines deformation learning with flow matching for anatomically consistent multi-organ image synthesis. We introduce a variational deformation network that learns a source-conditioned prior distribution over plausible anatomy-specific deformation fields, capturing the statistics of residual misalignment without requiring paired target inputs at inference, unlike existing registration-guided methods. This integration achieves efficient and stable image generation with substantially fewer integration steps compared to diffusion-based counterparts while preserving anatomical fidelity. The proposed framework is trained and evaluated on a large-scale dataset of 873 patients across five anatomical regions: abdomen, brain, pelvis, head-and-neck, and thorax. Extensive quantitative and qualitative evaluations demonstrate consistent performance in synthesis accuracy across all anatomical regions, while maintaining computational efficiency. Ablation studies further validate the efficacy of this approach, highlighting the individual contributions of learned deformations and flow matching over registration-guided and diffusion-based baselines.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/4912_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

https://github.com/pks716/MICCAI_2026

Link to the Dataset(s)

N/A

BibTex

@InProceedings{SinPee_VarDeFlow_MICCAI2026,
        author = { Singh, Peeyush Kumar AND Gulzar, Inam Ul Haq AND Singh, Sneha AND Nigam, Aditya AND Gupta, Pankaj},
        title = { { VarDeFlow: Variational Deformation Learning with Flow Matching for Multi-organ Cross-Modality Medical Image Synthesis } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16889},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    The authors propose a generative model which integrated variational deformation learning with flow matching to address residual spatial misalignment without requiring registered inputs at inference. The authors propose to initialize the flow matching network with a warped image, disentangling appearance from registration accuracy.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    • The idea of using a variational deformation learning as a prior for the flow matching network is interesting and seems novel.
    • The approach is more efficient during inference since it requires only 5 Euler steps as opposed to hundreds for diffusion-based approaches.
    • The method is evaluated on a large public dataset containing different modalities, multiple metrics are used for evaluation, an ablation study shows the importance of the deformation field estimation.
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
    • The visual results in Fig. 3, especially for the abdomen and thorax subsets, do not look convincing.
    • A more in-depth investigation on the disentanglement of geometric deformation and appearance shift would be interesting (How much deformation is actually left for the flow matching network?).
    • The synthesis is only done in one direction: MRI to CT. However, CT to MRI would also be interesting.
    • The presentation of fig. 3 could be nicer, the images look squished.
  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The idea of the paper is interesting and novel. A few issues as mentioned above could be improved.

  • Reviewer confidence

    Somewhat confident (2)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #2

  • Please describe the contribution of the paper

    This paper presents VarDeFlow, a flow-matching-based network for medical image synthesis. The proposed method jointly models deformation learning and image synthesis within a unified framework, and demonstrates improved performance compared with the benchmark methods. Moreover, relative to diffusion-based approaches, the flow-matching strategy requires fewer inference steps, suggesting a clear advantage in efficiency.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1.The method has been validated on five datasets covering different anatomical regions, such as the abdomen and brain, each exhibiting different magnitudes of deformation. This comprehensive evaluation supports the robustness and generalizability of the proposed network. 2.Experimental results show that the proposed network achieves superior performance to benchmark methods on all evaluation metrics, including MAE, SSIM, and PSNR, and offers improved efficiency compared with diffusion-based methods.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1.The paper fails to clarify whether the benchmark methods that are not registration-guided are performed with or without pre-registration, which makes the experimental setting insufficiently transparent. 2.In addition, the absence of a comparison between the proposed model with flow-matching method w/ external pre-registration limits the completeness of the ablation study evaluation.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (5) Accept — should be accepted, independent of rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper proposes a clear and strong method for medical image synthesis, which efficiently outperforms the benchmark methods and is supported by publicly available code. The minor limitations do not diminish the overall significance and strengths of the paper.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    As state above, the methodology is creative and the performance is also good enough. I think it is worth for accept



Review #3

  • Please describe the contribution of the paper

    The paper VarDeFlow addresses the challenge of spatial misalignment in cross-modality medical image synthesis by introducing a framework that combines a variational deformation network with a flow matching synthesis model. The variational deformation network learns a distribution over plausible deformation fields using only the source image at inference, which allows for effective alignment correction without requiring a target modality. Additionally, the model uses continuous normalizing flows via optimal transport to achieve high-quality synthesis in as few as five integration steps, offering a significant efficiency improvement over standard diffusion models. The authors also demonstrate the robustness of their method by training on a large-scale, multi-organ dataset encompassing 873 patients across five distinct anatomical regions including the abdomen, brain, pelvis, head and neck, and thorax using the SynthRad2023 and 2025 datasets.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1.The use of 873 patients from both SynthRAD2023 and 2025 grand challenges provides a robust and publicly verifiable dataset for training and evaluation.

    2.The formulation of a variational network to capture the statistics of residual misalignments is a theoretically sound approach to the registration-guided synthesis problem.

    3.The model offers high flexibility in controlling the inference process by allowing for the adjustment of the number of integration steps. The near-linear transport paths established through flow matching enables a faster inference than diffusion model, 3D-patch-based, and the complexity of the synthesis can be scaled based on specific clinical or computational requirements,

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1.The authors compare VarDeFlow against RegGAN (2021) and ResViT (2022). Given that this paper utilizes the SynthRad2023/2025 datasets, it is quite problematic that the authors do not compare against the established SOTA methods from those challenges, such as 3D patch-based nnU-Net or SwinUNETR. Comparing to an outdated 2D-based approach like RegGAN provides a misleading sense of superiority.

    2.While the pelvis results appear decent, the qualitative results in Figure 3 for the thorax and abdomen are alarming. There is a visible lack of structural fidelity in bone reconstruction and, more critically, significant artifacts within the lung parenchyma. This indicates a poor implementation of the synthesis module that fails to meet the standards set by top-tier challenge entries. Also, the layout of Figure 3 is not adapted to standard document widths, making it difficult to read and further highlighting the poor quality of the visual reconstructions.

    3.The evaluation is restricted to intensity-based metrics (PSNR, SSIM, MAE). For radiotherapy-focused data, the authors must include segmentation-based metrics or dosimetric validation (e.g., Gamma index) to prove anatomical plausibility.

    4.The authors should have compared their results directly against the online challenge leaderboards (which does provide segmentation and dose-based metrics automatically). Showing that their variational approach can enhance a standard baseline on deformed volumes would have been a much stronger scientific proof of concept.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (2) Reject — should be rejected, independent of rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    While the paper presents a technically promising idea in the Variational Deformation Network, the empirical execution is highly flawed. The use of outdated baselines on a modern, publicly benchmarked dataset is a significant scientific oversight. Most importantly, the qualitative results are extremely poor in challenging regions like the lungs and abdomen, failing to reach the fidelity expected from recent SynthRAD submissions. The lack of clinical or segmentation-based metrics further prevents an assessment of the method’s anatomical plausibility. To improve the paper, the authors must benchmark against modern 3D leaderboards (nnU-Net/SwinUNETR), provide segmentation/ dosimetric validation, and address the severe artifacts in their synthesis output.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    I thank the authors for their response and for providing local benchmark numbers for nnU-Net and SwinUNETR. The theoretical integration of variational deformation modeling with flow matching is interesting and mathematically sound. However, the empirical execution and evaluation remain problematic. The authors provided baseline comparisons only for the low-deformation anatomies (Brain and Pelvis) while omitting them for the highly challenging Thorax and Abdomen regions, where their model shows alarming structural artifacts and lung parenchyma anomalies in Figure 3.Attributing these artifacts to a “squished layout” or physical MR signal limitations does not justify the lack of structural fidelity, which top-tier entries on the SynthRAD leaderboard usually manage to resolve using standard 3D approaches. Furthermore, it remains unclear why the authors did not submit their method to the public challenge platform. An official submission would have automatically yielded independent, unbiased evaluation scores for both dosimetry and segmentation. Even if the challenge’s internal preprocessing differs slightly from the authors’ local setup, it would have provided crucial validation that this framework can achieve clinically relevant performance under an objective, standardized benchmark.



Author Feedback

We thank all reviewers for their insightful comments, recognizing novelty of our variational deformation prior, efficiency over diffusion baselines, and breadth of our 873-patient multi-anatomy evaluation with code release. We address all concerns, with committed revisions in camera-ready to strengthen our work. [R1]-W1,W4: Visual Quality of Fig-3: Abdomen/thorax are intrinsically challenging anatomy due to motion, pulsation, and soft-tissue deformation, resulting in visual imperfections which are shared across reported methods. Nonetheless, we achieve statistically significant improvements over baselines on all three metrics (p<0.05): MAE and SSIM for thorax (113.65 HU/0.822) and abdomen (103.11 HU/0.836). Error maps Fig 3 consistently show larger blue (low-error) regions for our method. Compressed Fig.3 layout (5 anatomies×6 methods) unfortunately obscures these differences; it is redesigned with native aspect-ratio and zoomed boundaries. [R1]-W2: Deformation-Appearance Disentanglement Analysis: Learned scalar (α) in our tanh-bounded displacement formula [ϕ = α·tanh(Dψ(z))] (Sec 2.1) directly quantifies it. α converges to 0.03-0.05 across anatomies, constraining max. displacement to 2.9-4.8mm for 96³ patches at 1mm³ resolution which is precisely the scale of residual misalignment after pre-registration (Table 1). The Variational Deformation Network (VDN) handles local mm-scale geometric correction, flow matching then performs contrast transformation on xwarped. As suggested, per-anatomy deformation analysis added in Sec 3.[R1]-W3: Bidirectional Synthesis: We prioritized MRI-CT as it is a prime use case for radiotherapy planning [22]. Our framework is symmetric; VDN operates on source images regardless of modality. Flow-matching model only requires informed prior that can be constructed for either direction. CT-MRI is added to Future Work as a direct extension. [R2]-W1: Experimental Transparency: All baselines in Table 2 are trained and evaluated on identical Elastix[10] pre-registered data (Sec 3). ResViT and Diffusion models receive same pre-registered inputs as RegGAN and RegConDIS, added explicitly in Sec 3.[R2]-W2: Ablation Completeness: This comparison is implicitly made in Table 3 as Flow Matching w/o deformation (N=0) [row 2], uses only Elastix pre-registered xsource as informed prior, confirming FM+deformation [rows 3-5] outperforms pre-registration alone. [R4]-W1: Baseline Selection: nnUNet/SwinUNETR are top performing methods on SynthRAD leaderboard, but they are deterministic regressors structurally unable to model residual post-registration misalignment hence we selected registration-guided (RegGAN, RegConDIS). Nonetheless, we train both under identical splits: nnUNet- MAE/PSNR/SSIM of 60.26/28.74/0.884(brain) and 49.96/29.33/0.881(pelvis); SwinUNETR- 61.33/28.44/0.878(brain) and 50.23/29.11/0.876(pelvis). Ours consistently outperforms leaderboard-deployed methods (p<0.05): 58.48/29.37/0.896(brain) and 47.76/30.51/0.902(pelvis), addressing leaderboard concern (added in Table 2). Clarification: We trained 3D-RegGAN under identical protocol (Sec 3). [R4]-W2,W4: Synthesis Quality, Fig.3 layout & Leaderboard: (a) Lung parenchyma synthesis is theoretically limited for all methods due to near-zero MR signal in air-filled lungs. Thorax SSIM clusters at 0.80-0.82 for all baselines with ours highest (0.822); MAE is lowest across all methods (Table 2) with only 5 integration steps, addressing the weak synthesis concern (b) Fig.3 layout has been redesigned ([R1]-W1,W4) (c) Leaderboard comparison is infeasible as SynthRAD evaluates on non-public held-out test set with challenge-specific preprocessing. We compare against nnU-Net/SwinUNETR ([R4]-W1). [R4]-W3: Gamma Index: Good suggestion but it requires RTPLAN/RTDOSE data, unavailable in public SynthRAD release and performed via organizers’ internal pipeline. Segmentation suggestion is insightful; DSC via pre-trained model is valuable extension, added in future work.




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    This paper tackles an important and challenging problem by proposing the novel and conceptually appealing idea of combining variational deformation modeling with flow matching. The efficiency gains over diffusion-based models (requiring only a few integration steps) is a key advantage. Moreover, the thorough experimental evaluations (multiple datasets, anatomical regions, and metrics) suggest robustness and generalizability, and provide insightful ablations. The paper is clearly written and reproducibility is strengthened by evaluating on public datasets and releasing the code.

    However, concerns were raised about potentially outdated or suboptimal baselines, with missing comparisons to stronger and more recent methods (e.g., modern 3D architectures or challenge leaderboards). Moreover, the poor results shown in Figure 3 undermine the quantitative claims and need more explanations.

    Further concerns include the lack of clarity regarding pre-registration settings, missing ablations against simpler alternatives, the lack of task-driven metrics, and figure readability.

    Overall, this paper presents a promising and potentially impactful idea, but I would like the authors to comment on points brought up by the reviewers, especially about the experimental setup and results.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    I thank the authors and reviewers for the fruitful discussion. Authors satisfyingly clarified concerns about fairness in pre-registration settings for all competing methods, brought new insights about a missing ablation which was already included, and promised to improve figure readability. They also explained why clinical metrics are hard to add right now.



    However, substantial concerns remain, notably about the experimental set-up, which does not include strong baselines for all experiments. The authors have added strong comparisons in 2 scenarios out of 4 (brain and pelvis), but failed to provide a satisfying explanations as to why they haven’t also done it for the 2 other scenarios (thorax and abdomen, which are also the most complex to handle). Moreover, the author refused to submit to challenge leaderboards, which would have provided a clean unbiased evaluation including required clinical metrics.

    This places this paper as borderline accept/reject. All things considered, I lean towards acceptance given the strong methodological aspect of this paper that introduces a novel and interesting framework. However, this acceptance is conditional on the authors adding what they promised in the camera ready (notably the new baselines and the refined fig 3).



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    This paper presents an approach for image synethesis using flow matching and variational deformation modeling. There were some initial concerns regarding the quality of the obtained results and the evaluation strategy all three reviewer recommend acceptance after the rebuttal.



Meta-review #3

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The paper received uniformly positive recommendations following the rebuttal. Reviewer 3 revised their assessment from Reject to Accept, Reviewer 1 also upgraded their score to Accept, and Reviewer 2 maintained an Accept recommendation. Overall, the reviewers agreed that the proposed method is interesting and that its empirical performance is strong and appropriate for the medical image synthesis setting. In light of the post-rebuttal discussion and the clear reviewer consensus, I recommend acceptance.

    For the camera-ready version, the authors are encouraged to further improve the clarity of the presentation and to address the reviewers’ constructive suggestions where appropriate.



back to top