Abstract

Developing robust artificial intelligence models for 4D (3D + time) medical imaging is constrained by limited annotated data, inter-device domain shifts, and privacy restrictions. To address this, we propose a 4D controllable generative framework for anatomically consistent data augmentation. The model generates both volumetric intensity sequences and their corresponding, inherently aligned segmentation masks. A semi-supervised variational autoencoder learns a compact latent representation of anatomical volumes while jointly predicting aligned segmentation masks in a unified framework. Anatomical structure is then disentangled from temporal dynamics through a cascaded latent diffusion model (LDM). A static LDM generates subject-specific anatomy conditioned on clinical priors (diagnosis and volumes measures) and a subsequent motion LDM estimates residual latent motions, ensuring strict temporal coherence across the 4D sequence. The proposed approach was evaluated on cine cardiac MRI as a representative 4D imaging application. Experiments across multiple datasets demonstrate high controllability of static anatomy (Pearson r > 0.8) and strong temporal coherence (FVD = 288.08). In cross-vendor generalization experiments, augmenting training sets with synthetic 4D sequences significantly improves downstream segmentation performance. Using nnU-Net, the proposed augmentation strategy improves the average Dice score by 1.4% and reduces the Hausdorff Distance by 3.0mm compared to training on real data alone, for the left ventricle, Dice improves by 2.8% with a 5.4mm reduction in boundary error. Overall, this framework provides a scalable and controllable solution for 4D medical image synthesis, supporting the development of more robust models with limited annotations and cross-vendor variability. Code available on https://github.com/cyiheng/4DCardiacMRISynthesis.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/3003_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: https://papers.miccai.org/miccai-2026/supp/3003_supp.zip

Link to the Code Repository

https://github.com/cyiheng/4DCardiacMRISynthesis

Link to the Dataset(s)

ACDC dataset: https://www.creatis.insa-lyon.fr/Challenge/acdc/databases.html Kaggle dataset: https://www.kaggle.com/c/second-annual-data-science-bowl/data M&Ms dataset: https://www.ub.edu/mnms/ M&Ms2 dataset: https://www.ub.edu/mnms-2/

BibTex

@InProceedings{CaoYih_AnatomyGuided_MICCAI2026,
        author = { Cao, Yiheng AND Andrade-Miranda, Gustavo AND Zhang, Jiatian AND Zhao, Lingxiao AND Gao, Xin},
        title = { { Anatomy-Guided Residual Motion Diffusion for Controllable 4D Cardiac MRI Synthesis } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16890},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    This paper presents a controllable framework for 4D cine cardiac MRI synthesis under limited annotation and domain shift. The method first learns a semi-supervised latent representation with a 3D VAE-GAN that jointly reconstructs images and predicts segmentation masks, so that generation takes place in a semantically informed latent space. The 4D synthesis problem is then decomposed into two stages: a static latent diffusion model generates ED anatomy conditioned on clinical priors, and a motion latent diffusion model generates ED-referenced residual motions over time. The resulting latent sequence is decoded into paired 4D image and segmentation outputs. The paper further evaluates the practical value of the synthetic data for improving downstream cardiac segmentation, especially under cross-vendor generalization settings.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    The paper addresses a relevant and challenging problem, namely controllable 4D cardiac MRI synthesis under limited annotation and domain shift. The proposed anatomy-motion decomposition is sensible. Modeling the sequence as static ED anatomy plus residual motion gives the framework a clear structure and makes the temporal generation process easier to interpret. The semi-supervised latent space is practically useful, since it enables direct generation of paired image-mask data without requiring dense labels for all subjects. The experimental evaluation goes beyond generative metrics and includes downstream segmentation on multiple backbones and external cross-vendor datasets. The gains on M&Ms and M&Ms2, especially for LV Dice and HD, are encouraging and support the practical value of the approach.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    The overall novelty is somewhat limited. The framework is reasonable and well organized, but many of its main ingredients are close to recent work on cardiac MRI video generation, conditional 4D cardiac synthesis, and spatiotemporal disentanglement. In particular, the paper would benefit from a sharper discussion of how the proposed ED-referenced residual motion formulation differs from prior approaches such as GANcMRI, TexDC, and recent disentanglement-based 4D cardiac generation methods. The paper does not provide direct comparison with the most relevant prior 4D cardiac generation methods. This makes it harder to judge how much is gained from the proposed design relative to existing approaches. The ablation study is too limited for a method paper of this type. Several components appear central to the claims, but their individual contributions are not isolated. At minimum, I would expect ablations on: removing the segmentation-aware latent supervision; removing the clinical conditioning. These experiments would make it much clearer which design choices are actually responsible for the downstream gains.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    This paper tackles a relevant and difficult problem, and the proposed framework is well motivated. The anatomy-motion decomposition is sensible, and the use of a semi-supervised latent space for paired image-mask synthesis is practically useful. The strongest part of the paper is the downstream validation on external cross-vendor datasets, which suggests that the synthetic data has real value for improving segmentation robustness. My main reservations are about novelty and experimental depth. The overall contribution is meaningful but not highly novel, and the distinction from the closest recent work is not fully clear. The paper also lacks direct comparison with the most relevant prior 4D cardiac generation methods, and the ablation study is limited. For these reasons, I see the paper as slightly above the acceptance threshold, but still clearly borderline.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The authors have responsed my comments.



Review #2

  • Please describe the contribution of the paper

    This paper proposes a cascaded generative framework for 4D cardiac MRI synthesis. A semi-supervised VAE is first used to project images into a latent representation. A static latent diffusion model then generates ED latent base anatomy conditioned on clinical priors. A motion predictor estimates residual offset relative to the ED frame. A subsequent motion diffusion model synthesizes the full 4D sequence conditioned on the generated anatomy, clinical information, and estimated motion residuals. The resulting synthetic sequences are used as data augmentation, yielding improved downstream segmentation performance.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1.The paper addresses the challenge about limited annotated data and cross-vendor generalization. The overall framework is clearly presented and the motivation is easy to follow. 2.The augmentation strategy is evaluated across multiple datasets using nnU-Net and other baseline models, with consistent improvements in both Dice score and Hausdorff Distance.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1.The proposed cascaded LDM disentangle static anatomy and temporal motion. This resembles appearance-motion decomposition strategies well-established in video generation literature. The manuscript lacks a clear discussion of what technically distinguishes this approach from such prior works, and why existing video generation paradigms cannot be directly applied to 4D medical imaging. As presented, the contribution reads primarily as a domain transfer of known techniques rather than a methodological advancement. 2.The evaluation compares only against a no-augmentation baseline in downstream segmentation, while the generative model itself is not benchmarked against state-of-the-art synthesis methods. Specifically, FID and FVD scores are reported in isolation without comparison to existing 4D cardiac generative models such as cine cardiac GANs or optical flow conditioned models. This makes it difficult to determine whether the observed performance gains stem from the specific design of the proposed framework or simply from the inclusion of any synthetic augmentation. 3.Tables 1 and 2 contain abbreviations that are not formally defined, which reduces the readability. 4.Key design choices lack ablation study. For example, the necessity of the semi-supervised VAE over supervised or unsupervised alternatives, the two-stage cascaded LDM versus a single unified diffusion model, and the actual contribution of clinical prior conditioning each require dedicated analysis. 5.The proposed framework generates motion via a residual latent motion LDM, but no comparison is provided against alternative motion generation strategies within the same pipeline, such as simple frame-by-frame subtraction in latent space or other explicit motion extraction methods. 6.Fig. 2 presents synthetic volumes without any side-by-side comparison against real counterparts. Including representative real images as a reference would allow readers to more intuitively assess the visual fidelity and anatomical plausibility of the generated outputs. 7.The font size used in the tables appears smaller than the required 8pt minimum. This should be corrected to comply with submission formatting guidelines.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The major factors are the insufficient novelty justification relative to existing video generation methods, the lack of comparison against state-of-the-art generative baselines, and the missing ablation studies that would validate individual component contributions.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Reject

  • [Post rebuttal] Please justify your final decision from above.

    While the rebuttal clarifies some design motivations and commits to formatting fixes, the core weaknesses remain unaddressed.



Review #3

  • Please describe the contribution of the paper

    This paper introduces a set of models for conditional 4D cardiac MRI synthesis. Generative models can address the problem of cMRI data and label scarcity and further promote deep learning efforts. The proposed framework combines three components: a semi-supervised variational autoencoder that learns a compressed latent representation of 4D volumes guided by a segmentation objective; a static diffusion model that synthesizes a single reference frame; and a residual latent motion model that predicts the remaining frames. Together, these modules generate realistic 4D cMRIs alongside spatially aligned segmentation masks. The approach is evaluated using FID and FVD to assess image and video quality, Pearson correlation between the specified and measured conditioning variables, and a downstream segmentation task in which augmenting training data with synthetic samples yields measurable performance gains.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    • 4D cMRIs are hard to model, a suite of models including VAE and 2 LDMs represent a robust and creative approach.
    • Strong correlation between clinical volume priors and measured synthetic volumes shows the model’s ability to do attribute manipulation
    • Results demonstrate that synthetic data generated using this model can be effective augmentation method for training downstream deep learning models
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
    • The paper cites several directly related 4D CMR synthesis approaches (4D CardioSynt [5], TexDC[15, GANcMRI [18], Temporal Differential Fields [20]) but does not compare against any of them in FID, FVD, or downstream segmentation. The reported FID of 72.21 and FVD of 288.08 are presented in isolation, making it impossible to contextualize the method’s generative quality relative to previously developed models. For example, GANcMRI reports FID of 92.58 and FVD of 283.53, are these metrics directly comparable? A comparison against at least one of these models (desirably the best out of them) would significantly strengthen this paper.
    • Very small dataset sizes, specifically the test and a held-out subset consisted of 50 patients only.
  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    This paper presents a well-developed method for modeling 3D+time medical data, supported by a comprehensive set of experiments demonstrating strong visual fidelity, effective conditional generation, and utility as a data augmentation technique. The experimental validation is thorough and convincing. The approach also shows clear potential for extension to other imaging modalities. However, the rebuttal should provide the comparison with previously developed cMRI generation models.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The rebuttal addressed my main concerns. The authors provided a textual comparison with prior generative cMRI models and clarified the limitations of the small dataset size.



Author Feedback

We sincerely thank the reviewers for their valuable feedback. Our responses are included below:

Q: Novelty relative to prior 4D synthesis and video generation [R1,R2]

A: Our framework is not a direct transfer of generic video generation; it is designed specifically for controllable 4D medical augmentation under sparse annotations:

  • Joint image-mask synthesis via semi-supervision: Existing 4D MRI approaches generally synthesize intensity sequences only, requiring external pseudo-labeling for downstream training. In contrast, our semi-supervised VAE directly generates anatomically aligned 3D image-mask pairs from partially labeled and unlabeled data, enabling controllable augmentation without pseudo-label drift.

  • ED-referenced residual motion: Unlike conventional video models and recent cardiac frameworks that jointly evolve anatomy and dynamics or predict explicit 3D deformation fields, we model temporal evolution strictly as stochastic latent residuals (zt=zED+mt). This explicitly anchors the entire 4D sequence to a single subject-specific 3D ED anatomy, transforming generation into an anatomy-preserving latent deformation. This design aims to preserve topological correspondence across time, reduce temporal drift, and facilitate memory-efficient, high-resolution 4D synthesis.

Q: Lack of direct comparisons and interpretation of FID/FVD [R1,R2,R3]

A: We agree that standardized benchmarking against prior 4D cMRI generation methods is important. However, direct reproduction is challenging because several recent approaches rely on unavailable implementations, restricted datasets, or substantially different 2D/3D settings. In addition, prior methods generally generate intensity sequences only, requiring external pseudo-labeling for downstream evaluation, which unfairly entangles the generator’s quality with the pseudo-labeler’s errors. Regarding R3’s question on GANcMRI, cross-paper FID/FVD comparisons are only strictly valid under identical datasets and protocols. Additionally, GANcMRI is 2D-based whereas our framework models full 3D volumes. While such cross-setting comparisons should be interpreted cautiously, TexDC reported FID=56.13 and FVD=756.92 with similar public data-mix, whereas our framework achieves FID=72.21 and FVD=288.08, suggesting stronger temporal coherence.

Q: Justification for design choices & ablations [R1,R2]

A: We agree that ablation studies would provide further insight, and this will be addressed in future work. Below, we clarify the motivation behind the current design choices:

  • Semi-supervised VAE: A fully supervised setting reduces usable data (no label for DSB), while unsupervised learning fails to preserve the explicit image-mask correspondence required for paired synthesis. The semi-supervised design therefore enables the use of unlabeled cine data while maintaining anatomically aligned image–mask generation.

  • Two-stage decomposition: Separating anatomy and motion is intended to promote anatomical stability and temporal consistency. This factorization reduces the learning burden compared to a unified model by isolating static and dynamic complexities, which is qualitatively consistent with the low FVD observed in our experiments.

  • Clinical conditioning: Conditioning enables controllable oversampling of rare or clinically relevant phenotypes, which is a primary intended application. A natural way to further assess this capability would be to evaluate segmentation performance using pathology-enriched synthetic training data on underrepresented disease subgroups.

Q: Dataset size and formatting [R2,R3]

A: Limited annotated 4D MRI data is precisely the motivation for this work. To ensure robustness beyond the 50 internal test cases, we validated on two entirely unseen, multi-vendor external datasets. Evaluating on nearly 300 unseen multi-vendor external cases demonstrates strong cross-vendor generalizability. Finally, we will correct the reported formatting and readability issues.




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    This paper proposes a controllable 4D cardiac MRI synthesis framework that decomposes generation into static ED anatomy and residual motion via cascaded latent diffusion models. Reviewers find the proposed method is somewhat useful and the downstream cross-vendor segmentation gains encouraging. However, they consider the methodological novelty limited relative to existing video generation and 4D cardiac synthesis methods. They also point out missing comparisons with prior 4D cMRI generative models and insufficient ablations. Given these mixed assessments, the paper moves to rebuttal for the authors to address the concerns.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The rebuttal clarified the key design distinctions and contextualised the FID/FVD results, leading two reviewers to accept post-rebuttal. The remaining reject reflects concerns about novelty and missing ablations that are real limitations but do not outweigh the majority consensus given the method’s clear clinical utility and strong external validation. The recommendation is therefore Accept.



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    This paper proposes an anatomy-guided residual motion diffusion framework for controllable 4D cardiac MRI synthesis, addressing challenges such as limited annotated data, inter-device domain shifts, through a semi-supervised VAE and cascaded LDMs. Main concerns focus on its unclear distinction from previous work, performance comparison, and insufficient ablation studies. The authors’ response explain these clearly or reasonably. Therefore, I think when revised accordingly, the paper can be accepted.



Meta-review #3

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Reject

  • Please justify your recommendation.

    The paper addresses an important problem and shows promising downstream segmentation gains with synthetic 4D cardiac MRI data. However, the main concerns remain after rebuttal, especially the lack of direct comparison with relevant 4D cardiac synthesis methods and the absence of ablations for the key components of the proposed framework. Since these issues affect the strength of the methodological claims and cannot be resolved by minor camera-ready changes, the paper is not recommended for acceptance.



back to top