Abstract

Conditional brain templates and atlases are essential for cohort-specific groupwise registration and for monitoring age-related or disease-imposed changes by quantifying anatomical and regional differences. However, most existing methods construct templates and atlases independently, leading to suboptimal quality of both. Besides, atlas construction often requires individual parcellation labelmaps, which are rarely available in practice. To address these limitations, we propose a unified framework that jointly generates conditional templates and atlases, requiring only an arbitrary pair of publicly available template and atlas. The model learns conditional prototypes in a shared latent space and maps them to templates and atlases via task-specific decoders, enabling effective interaction between anatomical and regional features. To further overcome the absence of individual labelmaps, we adopt a one-shot strategy that propagates the public atlas to individual subjects to build pseudo labelmaps, supervising the parcellation decoder. By optimizing prototypes across both latent and volumetric spaces over the entire population, the decoded templates maintain anatomical fidelity comparable to real MRIs, while both templates and atlases embody cohort-specific characteristics and model smooth variations across conditions. Comprehensive evaluations across four public datasets using eight multifaceted metrics demonstrate that our framework achieves favorable overall performance compared to recent state-of-the-art methods.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/3218_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

https://github.com/LSYLAB/LatentAtlas

Link to the Dataset(s)

Preprocessed OASIS dataset: https://github.com/adalca/medical-datasets/blob/master/neurite-oasis.md IXI dataset: https://brain-development.org/ixi-dataset/ Mindboggle101 dataset: https://dataverse.harvard.edu/dataset.xhtml?persistentId=doi:10.7910/DVN/HMQKCK IBSR18 dataset: https://www.nitrc.org/projects/ibsr/

BibTex

@InProceedings{ZhaJic_Learning_MICCAI2026,
        author = { Zhang, Jichang AND Che, Tongtong AND Li, Feng AND Zhang, Lin AND Wang, Chengbo AND Wang, Xiuying AND Li, Shuyu},
        title = { { Learning Prototypes for Unsupervised Joint Generation of Conditional Templates and Atlases } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16889},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    The paper proposes LatentAtlas, an unsupervised framework for jointly generating age-conditional brain templates and atlases from a single public (template, atlas) pair. A multi-decoder VAE learns a shared latent space mapped to images, labelmaps, and deformation fields; a FiLM-conditioned generator produces prototypes optimized in both latent and volumetric spaces. Pseudo-labels for the parcellation decoder are obtained by warping the public atlas to individual subjects via the registration decoder. Evaluated on OASIS (train) and IXI/MIND/ISBR (test) against AtlasMorph, DiffDef, and InGroup, across eight metrics.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    • Well-motivated problem and coherent formulation. Joint template–atlas generation with only a public (template, atlas) pair is a realistic, practically useful setting, and the latent-prototype formulation cleanly unifies the three tasks (reconstruction, parcellation, registration).
    • Strong OOD results. On IXI, Ours(ISBR) improves Mov (488 -> 456) and Norm (762 -> 722) over all baselines with p < 0.01; on MIND, Ours(MIND) improves Norm (804 -> 759) with significance. Topology (A. Topo. ) and transition smoothness (A. Trans. ) are best on both public-prior settings.
    • Very comprehensive evaluation. Four datasets, two public priors, eight metrics, paired t-tests, and a released anonymous code repository (more than typical MICCAI papers). Furthermore, the code is of very high quality.
    • Thoughtful design details. Gradient detachment between prototype losses and the shared encoder/decoders, gated normalization to handle intensity mismatch between individual MRIs and the public template, and dual-space (latent + volumetric) prototype regularization are all well-reasoned.
    • Informative Fig. 3 analyses. Ventricular trajectory consistency across MIND and ISBR priors, and the UMAP visualization showing prototypes forming an age-ordered continuum embedded within individual features, support the method’s claims.
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
    • No ablation study. The method has many components (parcellation decoder, latent-space alignment with L2 + cosine, volumetric centrality, gradient detachments, FiLM generator), yet none are ablated (most probably due to space constraints). The contribution of each cannot be assessed from the current results.
    • “Superior holistic performance” is slightly overstated. On OASIS (source), InGroup wins Mov (236 vs 246/251) with significance; MIND DSC ties DiffDef (0.52); ISBR DSC (0.66 vs 0.65 for AtlasMorph/InGroup) is within noise. “Outperforms all baselines on OOD” is accurate on Mov/Norm/LPIPS but not DSC.
    • T. Trans. on templates is worse than AtlasMorph on both MIND (. 33 vs . 17) and ISBR (. 29 vs . 17). Not acknowledged in the text, which emphasizes the A. Trans. (atlas) win.
    • Topology metric interpretation unclear. A. Topo. differs by ~75× between MIND (-3260 baseline) and ISBR (-44 baseline), only partly explained by class count (95 vs 34). All methods, including Ours, produce highly non-trivial topology; “simpler” is relative, not absolute.
    • No pseudo-label quality analysis. The parcellation decoder is supervised by warped public atlases whose quality depends entirely on registration accuracy; no analysis of label noise or its effect is provided.
    • OASIS “source” comparison is confounded. InGroup constructs templates from age-nearest OASIS training neighbors, so comparing on OASIS favors it by design. The OOD comparisons are the meaningful ones.
    • T. Sharp. differences within noise. 0.90 vs 0.91 vs 0.92 with stds 0.01-0.02.The “sharper templates” claim is weakly supported.
    • Minor writing issue. “since they requires”
  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (5) Accept — should be accepted, independent of rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    This submission is clearly above the typical MICCAI bar: the problem is well-motivated, the method is coherent and thoughtfully engineered, the evaluation is multi-dataset and multi-metric with statistical testing, and the code is released. The absence of an ablation study is likely a space constraint given the 8-page limit and the breadth of the evaluation already included. Several textual claims slightly overstate mixed quantitative results and should be softened. Overall, a solid accept.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #2

  • Please describe the contribution of the paper

    The paper addresses an important and underexplored problem in neuroimaging, namely the unsupervised joint generation of conditional brain templates and atlases. The motivation is well grounded: existing methods typically construct conditional templates and atlases separately, even though the two are semantically coupled and should ideally reflect consistent anatomical and regional structure.

    The main contribution is the proposed LatentAtlas framework, which reformulates conditional template and atlas construction as conditional prototype learning in a shared latent space. The method combines a multi-decoder VAE backbone with task-specific reconstruction, parcellation, and registration decoders, allowing the framework to jointly model anatomical appearance, regional structure, and deformation. A key practical enabler of the framework is the one-shot pseudo-labeling strategy, which uses a warped public atlas to provide approximate supervision for the parcellation branch in the absence of individual labelmaps.

    A second contribution is the conditional prototype generator, which produces cohort-specific latent representations from scalar conditions such as age and regularizes them in both latent and volumetric spaces. This design aims to ensure that the generated templates and atlases are not only anatomically plausible but also representative of condition-specific population characteristics.

    The paper further provides a relatively comprehensive empirical evaluation across four public datasets using multiple metrics that assess template fidelity, atlas topology, transition smoothness, and registration quality. The reported results demonstrate that the proposed joint framework yields consistent improvements over prior methods, with particularly notable gains in atlas topology and in the recovery of age-related trajectories from conditional atlases.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    Strength 1: Well-motivated and reasonably novel problem formulation. The paper is built around a clear and important observation: conditional templates and atlases are intrinsically coupled, yet most prior methods construct them separately. Reformulating their joint construction as conditional prototype learning in a shared latent space is a conceptually clean and original design choice. The key contribution is less about any single architectural component and more about the unified formulation that allows anatomical appearance and regional structure to be modeled together rather than in a sequential or post-hoc manner.

    Strength 2: Relatively comprehensive and carefully designed evaluation. The experimental section is broad in both datasets and evaluation criteria. The paper evaluates four datasets and uses multiple metrics that cover several distinct aspects of quality, including template sharpness, transition smoothness, atlas topology, cohort centrality, perceptual fidelity, and registration accuracy. This provides a more complete picture than a narrowly metric-driven evaluation. The inclusion of both in-distribution and out-of-distribution settings also strengthens the case that the method is learning cohort-level structure rather than overfitting to a single dataset. The use of statistical testing is another positive aspect of the empirical study.

    Strength 3: A practical design choice for label-scarce settings. The one-shot pseudo-labeling strategy is practically meaningful because subject-level parcellation labels are often unavailable in realistic neuroimaging scenarios. Using registration-derived deformation fields to propagate a public atlas and provide approximate supervision for the parcellation branch is a reasonable and useful engineering solution. This design substantially broadens the applicability of the framework beyond settings where dense individual annotations are available.

    Strength 4: Reproducibility is supported by code release. The paper provides an anonymous code repository, which is a meaningful strength for a methodologically complex paper.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    Weakness 1: The generality of the conditioning framework is not fully demonstrated. The paper presents LatentAtlas as a general framework for conditional template and atlas generation, but all experiments are conducted using age as the only conditioning variable. Age is a relatively smooth and readily available scalar factor, and it is therefore a favorable test case. It remains unclear how well the framework would extend to other clinically relevant conditions, such as disease status, sex, or multi-factor conditioning. This does not invalidate the current results, but it does limit the strength of the broader generalization claims made in the paper.

    Weakness 2: The paper lacks a systematic ablation study. The proposed framework contains several interacting design choices, including the multi-decoder VAE, the pseudo-labeling mechanism, the latent alignment objective, the volumetric centrality regularization, and the gradient detachment scheme. However, the paper does not isolate the contribution of these individual components through ablation experiments. As a result, while the overall method performs well, it is difficult to determine which design choices are most responsible for the reported gains. This is an important limitation for a multi-component methods paper.

    Weakness 3: The evidence based on ISBR is limited by the small sample size. The ISBR test set contains only 18 subjects with annotated labelmaps, making it the smallest and noisiest basis for the DSC analysis. Since DSC is only reported on MIND and ISBR, and ISBR is particularly small, the conclusions supported by this metric are somewhat constrained. In particular, the reported standard deviations on ISBR are not negligible relative to the mean differences between methods. A more cautious discussion of this limitation would strengthen the experimental section.

    Weakness 4: The fairness of the adapted AtlasMorph comparison could be better justified. The paper compares against an adapted version of AtlasMorph with FiLM conditioning layers inserted, while using originally reported hyperparameters. This is a reasonable baseline construction choice, but it is not fully clear whether the adapted version was sufficiently validated or tuned under the current setting. Since the quality of the comparison partly depends on this adaptation, the paper would benefit from a clearer justification of how fairness was ensured.

    Weakness 5: The gradient flow design is important but insufficiently analyzed. The gradient flow management strategy appears to be an important part of the optimization design, especially given the number of coupled modules and losses in the framework. However, the paper only describes the detachment rules at a conceptual level and does not empirically study how these choices affect training stability or final performance. This makes it harder to assess whether the proposed optimization design is essential, or simply one workable implementation choice.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    Comment 1: Demonstrating that the framework can generalize beyond age conditioning would substantially strengthen the paper. Even a brief additional experiment with a second conditioning variable, such as biological sex as a binary condition, would help show that the method is not specific to a smooth continuous scalar setting.

    Comment 2: A more systematic ablation study would greatly improve the paper. At minimum, it would be helpful to isolate the contribution of the pseudo-labeling strategy and to compare the proposed joint training scheme against a simpler sequential alternative in which template and atlas generation are optimized separately. This would make the source of the reported gains much clearer.

    Comment 3: For the ISBR-based DSC results, please clarify whether the observed differences between methods are statistically significant given the small sample size of 18 subjects. If statistical significance cannot be established, this limitation should be stated more explicitly in the paper.

    Comment 4: Please provide more detail on how the FiLM-adapted AtlasMorph baseline was validated to ensure a fair comparison. In particular, it would be helpful to clarify whether this adapted baseline was re-tuned under the current setting or used directly with the original hyperparameters after adding FiLM layers.

    Comment 5 (Minor): Some panels in Figures 2 and 3 appear to have limited resolution, which slightly reduces the readability of the qualitative comparisons. Higher-resolution figures would improve the visual presentation, especially for the atlas overlays in Figure 2B.

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper addresses a well-motivated and relatively underexplored problem, proposes a technically coherent framework, and provides a stronger-than-usual empirical evaluation for this line of work. In particular, the joint formulation of conditional template and atlas generation as prototype learning is a meaningful contribution, and the one-shot pseudo-labeling strategy addresses an important practical limitation in label-scarce neuroimaging settings. The availability of code, the use of four datasets, the broad set of evaluation metrics, and the inclusion of statistical testing all contribute positively to the overall credibility of the paper.

    The main reasons I do not score the paper more highly are the absence of a systematic ablation study and the limited evidence for conditioning generalizability. Because the framework contains several interacting components, the lack of ablation leaves it unclear which design choices are most responsible for the observed improvements. In addition, although age conditioning is a reasonable and relevant proof of concept, it does not fully support the broader claim that the framework is generally applicable to arbitrary conditioning variables.

    Overall, I view the paper as technically solid and experimentally stronger than average in this area, with strengths that outweigh its current limitations. For this reason, my assessment leans positive. That said, the missing ablation evidence and the narrow conditioning demonstration are real weaknesses, and a rebuttal that clarifies these points would further strengthen confidence in the final decision.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #3

  • Please describe the contribution of the paper

    The LatentAtlas, an unsupervised conditional atlas construction framework, requiring an arbitrary public template and atlas pair as prior. The framework construct conditional template and atlas in a unified latent space. LatentAtlas achieves state-of-the-art results across eight multifaceted metrics, including template sharpness, topological fidelity, and age-trend fitting.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1.The framework learns conditional prototypes in a shared latent space that are mapped to both templates and atlases via task-specific decoders. 2.It uses a one-shot strategy to propagate the public atlas to unlabeled individual subjects, creating pseudo-label maps that supervise the parcellation decoder. 3.The regional boundaries could provide guidance to sharpen anatomical details of templates.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1.If the initial registration between the public template and a target subject is poor that common in subjects with significant anatomical deviations, the resulting pseudo-labels will be noisy or inaccurate. The template and atlas generation tasks are coupled in a shared latent space, these label errors can backpropagate and negatively impact the anatomical fidelity and sharpness of the generated conditional templates. The quality of generated Atlas would also be affected. 2.The framework requires balancing multiple competing loss terms, including reconstruction, parcellation, registration, and KL divergence. It would be sensitivity to hyperparameter balancing in multi-task learning. It is important to add hyperparameter analysis.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The LatentAtlas, an novel unsupervised conditional atlas construction framework in latent space, requiring an arbitrary public template and atlas pair as prior. The pseudo-label maps provide the anatomical details without mannual annotation. The experiment result verify the effectiveness of LatenAtlas in multi aspects, such as template sharpness, topological fidelity, and age-trend fitting.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Author Feedback

We thank the reviewers and meta-reviewer for their positive assessment and constructive comments. We respond below to the main concerns and clarify several points.

(1) On generalization beyond age conditioning. R2 and the meta-review raise the concern that all experiments use age as the conditioning variable, and therefore question whether the framework can generalize to other conditions. Our method is not restricted to age. In LatentAtlas, the input condition c can be extended by increasing its dimensionality, e.g., by including categorical variables such as sex, imaging site, or disease status, in addition to continuous variables. We also agree that constructing templates and atlases under more fine-grained and comprehensive condition settings, and studying their downstream applications, is an important direction for future work.

(2) On pseudo-label noise. R2 and R3 raise concerns that registration-derived pseudo-labels may be noisy, and that such noise could affect template and atlas generation. We agree that pseudo-label noise may exist. This is also a common challenge in one-shot/few-shot learning and weakly supervised settings. Our pseudo-labeling strategy is motivated by prior findings (e.g., DeepAtlas, 2019) that, despite label noise, networks can benefit from large amounts of imperfect supervision rather than relying only on very few accurate labels.

(3) On the ISBR dataset and DSC interpretation. R1 and R2 point out that ISBR contains only 18 manually labeled subjects and that the DSC improvement on ISBR is not statistically strong. Since gold-standard parcellation labelmaps manually delineated by experts are rare in public neuroimaging datasets, ISBR is valuable but limited in scale. Its small sample size can introduce additional noise and statistical uncertainty, especially for DSC evaluation where mean differences between methods are small. Therefore, the ISBR DSC results should be interpreted cautiously and together with broader OOD results, including Mov/Norm/LPIPS and intrinsic template/atlas quality. In the Results section, we will make this limitation clear.

(4) On the AtlasMorph comparison. R2 questions whether the FiLM-adapted AtlasMorph baseline is fair. Since FiLM is only used as a conditioning mechanism and is not the core contribution of our method, we adopted AtlasMorph+FiLM to avoid letting this non-core design choice affect the fairness of the comparison. Moreover, prior work (e.g., AtlasGAN, 2021) has shown that adding FiLM to AtlasMorph can improve over the original AtlasMorph setting. Thus, we believe that AtlasMorph+FiLM is a stronger and fairer baseline than the original AtlasMorph.

(5) On the wording of quantitative claims. R1 and R2 note that some claims may be too broad because the results do not show uniform superiority on every metric. We agree and will soften the corresponding statements in the manuscript to make the descriptions more rigorous and precise.

(6) We will also address minor presentation issues, including correcting the grammar issue “since they requires” and improving the resolution of Fig. 2. We sincerely thank the reviewers and meta-reviewer for their insightful suggestions, which help us improve the current work and also point to valuable directions for future research.




Meta-Review

Meta-review #1

  • Your recommendation

    Provisional Accept

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    Reviewers agree the paper proposes a novel and well-motivated framework for joint conditional template-atlas generation, with strengths in coherent formulation, practical pseudo-labeling, and comprehensive multi-dataset evaluation showing strong OOD performance. They also notice limitations including lack of ablation studies, limited generalization evidence (only age conditioning), and potential issues from pseudo-label noise and hyperparameter sensitivity. Overall, the work is technically solid and above average.



back to top