Abstract

Reconstructing 3D CT volumes from bi-planar X-rays offers a low-dose, efficient alternative to CBCT for Adaptive Radiotherapy. However, the anatomical overlaps in X-rays and the ill-posed nature of this task often lead to spatial ambiguity and a loss of patient-specific details in existing methods. In this paper, we propose EA-LDM, an Explicit Alignment Latent Diffusion Model framework to address these challenges. First, we introduce an Explicit Alignment (EA) strategy, where an alignment network maps 2D X-rays directly into a pre-trained 3D CT latent space, utilizing volumetric semantic priors to resolve spatial ambiguity. Second, to recover patient-specific details, we propose Patient- Specific Optimization (PSO). Formulated through a bi-level optimization, PSO effectively incorporates priors from the pre-treatment planning CT while maintaining robustness to temporal anatomical changes. Experiments on a collected clinical dataset demonstrate that our method significantly outperforms state-of-the-art approaches in both structural fidelity and perceptual quality.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/0820_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

N/A

Link to the Dataset(s)

N/A

BibTex

@InProceedings{XiaJie_EALDM_MICCAI2026,
        author = { Xiao, Jiewen AND Wu, Shuqiong AND Chen, Le AND Feng, Bin AND Jin, Fu AND Fan, Xin},
        title = { { EA-LDM: Explicit Alignment Latent Diffusion Model for Patient-Specific CT Reconstruction from Bi-planar X-Rays } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16888},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    This paper aims to address the limitations of cone-beam computed tomography (CBCT) in intraoperative settings by proposing a method for CT reconstruction from biplanar X-ray images. The proposed approach is patient-specific, incorporating individualized optimization to improve reconstruction accuracy. Furthermore, it enables intraoperative registration with preoperative CT data through explicit alignment latnet diffusion model (EA LDM), facilitating better anatomical correspondence during surgery.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    This paper focuses on the registration of intraoperative X-ray images with preoperative CT by leveraging CT reconstruction from only two biplanar X-ray views (anteroposterior and lateral). The proposed method incorporates patient-specific modeling and optimization, enabling it to account for inter-patient variability as well as temporal anatomical changes that may occur between preoperative imaging and surgery. These features enhance its suitability for real-world clinical applications.

    The approach has been extensively validated on standard datasets and demonstrates competitive performance when compared with existing state-of-the-art methods in the literature.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    The paper remains overly generalized, and does not adequately address how the proposed framework can be adapted or customized to different anatomical regions. While the methodology is presented as broadly applicable, there is limited discussion on anatomy-specific constraints, variations, or the need for specialized modeling strategies.

    Although the RADCURE dataset (focused on oropharyngeal anatomy) is used for evaluation, the paper does not clearly explain how the algorithm is trained or tailored for this specific anatomical domain. In particular, it is unclear whether the model incorporates anatomy-aware priors or whether it relies purely on generic feature learning.

    Furthermore, the use of pretrained CT-based weights for processing X-ray inputs raises concerns. Given the inherent domain gap between CT volumes and X-ray projections, such pretrained representations may not yield meaningful latent features for accurate 2D-to-3D reconstruction. A more detailed justification or domain adaptation strategy is needed.

    A central challenge that is insufficiently addressed is the reconstruction of a 3D volume from only two biplanar X-ray images (AP and lateral). This is a highly ill-posed problem, and the paper does not convincingly demonstrate how the proposed method overcomes the associated ambiguities.

    Additionally, the availability of preoperative CT scans suggests that the problem could instead be formulated as a 2D–3D registration task. In such a scenario, intraoperative X-rays could be registered to the preoperative CT, allowing surgical planning to proceed directly on the CT data without requiring full 3D reconstruction. The paper does not provide a clear comparison or justification for choosing reconstruction over registration.

    In this context, the motivation for employing EA-LDM for 3D reconstruction needs to be more explicitly articulated. Specifically, the advantages of this approach should be clearly contrasted against conventional 2D–3D registration methods in terms of accuracy, robustness, computational efficiency, and clinical utility.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not provide sufficient information for reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper does not sufficiently justify the choice of 3D reconstruction over conventional 2D–3D registration approaches. In scenarios where preoperative CT is available, 2D–3D registration is a well-established and often more accurate solution compared to reconstructed 3D volumes from sparse X-ray views. The authors should clearly articulate the advantages of their reconstruction-based approach, particularly in terms of accuracy, robustness, and clinical relevance, in comparison to existing registration techniques.

    Additionally, the proposed architecture lacks a clear mechanism for adapting to the specific anatomical region of interest. Different anatomical structures exhibit distinct geometric, textural, and pathological characteristics, which typically necessitate customized modeling strategies. The paper does not discuss whether the model incorporates anatomy-specific priors, constraints, or training adaptations, raising concerns about its generalizability and clinical applicability.

    The experimental validation is limited to standardized datasets, which may not adequately reflect real intraoperative conditions. The transition from controlled datasets to real surgical environments introduces several challenges, including patient positioning variability, imaging noise, occlusions, and workflow constraints. The paper does not describe any operating room (OT) setup, workflow integration, or practical deployment considerations. For applications such as tumor resection or biopsy, it is essential to outline how intraoperative imaging would be acquired, processed, and utilized within the surgical pipeline.

    Furthermore, the comparative analysis is primarily restricted to diffusion-based and GAN-based methods. While these are relevant, the omission of Neural Radiance Field (NeRF)-based approaches is notable. NeRF-based methods have demonstrated strong capabilities in multi-view 3D reconstruction with reduced hallucination artifacts, and their inclusion would provide a more comprehensive and balanced evaluation of the proposed method.

    How the weighting parameters for the loss functions lambda 1 to lamda 5 is selected must be mentioned

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #2

  • Please describe the contribution of the paper

    This work introduces EA- LDM, a framework for patient-specific 3D CT reconstruction from biplanar X-rays. To address the spatial ambiguity in X-ray images, the work proposes Alignment Networks that map 2D X-rays into a pretrained 3D CT latent space, and a bi-level Patient-Specific Optimization (PSO) strategy that balances population-level generalization with patient-specific refinement. The proposed method outperforms state-of-the-art approaches in PSNR and SSIM.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1.Comprehensive ablations clearly demonstrate the individual contributions of Explicit alignment (EA) and PSO. The downstream segmentation evaluation using reconstructed CTs across anatomies (Section 3.4) is a particularly strong validation. 2.The bi-level PSO strategy is an interesting formulation where the lower-level optimization ensures generalization while the upper-level refines patient-specific details.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1.Reproducibility is a concern. The weighting hyperparameter in Eq. 3 is undefined, and the variables “x M” and “x N” in Fig. 2(a) are not explained. Combined with the absence of open-source code, replication would be difficult. 2.Evaluation is limited to DRR-based reconstructions, and validation on real biplanar X-ray images would be necessary to better assess the clinical feasibility of the proposed approach.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not provide sufficient information for reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    Although the proposed EA and PSO strategies are well motivated and clearly contribute to the improved performance, the underlying latent diffusion model appears similar to prior works such as DIFR3CT. It would strengthen the paper if the authors more clearly articulate how their latent diffusion model differs from existing approaches. In Section 3.4, the cohort size used for this evaluation is not specified, which makes it difficult to assess the robustness of the results. Additionally, the analysis relies primarily on qualitative visualizations so, including quantitative metrics would further strengthen the claims. Also, to improve reproducibility, the paper would benefit from clearly defining all hyperparameters or providing an open-source implementation of the work.

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper presents a well-structured and detailed approach to CT reconstruction from biplanar X-rays, leveraging explicit alignment (EA) networks and patient-specific optimization (PSO) as its core contributions. The ablation studies clearly demonstrate the individual impact of EA and PSO, showing the value of each component. However, reproducibility concerns and the lack of evaluation on real X-ray images are notable gaps. These limitations led me to the decision

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The authors have addressed reproducibility concerns by clarifying hyperparameters and ambiguous symbols. However, the architectural novelty of the Alignment Network (AN) still does not feel sufficiently differentiated. An ablation study specifically isolating the contribution of AN would strengthen the paper and provide more convincing evidence of its actual impact.



Review #3

  • Please describe the contribution of the paper

    This paper shows a technique to construct 3D views from biplanar X-rays. Explicit alignment is used to generate 3D reconstructions fro 2D X-rays. There are two steps here, a lower level optimisation, which generates a population level image, and a higher level optimisation which generates a patient specific optimisation.

    An alignment network is used to extract and align the embeddings generated from the two X-ray images. Position aware fusion is used to combine features at different scales. GAN, L1, and perceptual losses are used.

    PSO (Patient specific optimisation) is used for a few further steps to further improve the image quality. This is done with patient specific priors.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    The formulation as an LDM is interesting and scientifically sound.

    Implementation steps are given in detail with the hyper parameters. The data is from a public dataset and can be used to replicate the results.

    The comparison is against SOTA methods. Many full reference metrics are used. I am pleased to see that the inference time is also supplied.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    It is unclear how the X-rays in the alignment section are generated. Is it sampled directly from the CT? If so, are the scanner characteristics necessary to setup the problem?

    The entire section with the PSO confuses me. If it needs a pCT, one does not save the time from conducting a CT.

    The quantitive section compares with state of the art methods. I am not sure how this approach works. Are the SOTA methods also applied to generate 3D views from biplanar scans?

    To me, it is not clear what problem this approach solves. One can create a complete limited angle scan with LDCT and an AEC. This will probably yield comparable images to the presented approach.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (2) Reject — should be rejected, independent of rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    I do not see a use case for this method. Additionally I feel several details are unclear. A rebuttal will not be enough to satisfy my concerns about this work.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Reject

  • [Post rebuttal] Please justify your final decision from above.

    The authors’s comments don’t satisfy my core question i.e. what is the main use case of this work?

    I don’t agree with the points made about the time needed for limited angle scans being enough to cause issues - unless custom hardware has been designed, there will also be a time differential in acquiring biplanar X-rays.



Author Feedback

We thank all reviewers for their insightful comments.

Adaptation to Specific Anatomical Structures [R1-1/2] EA-LDM is not limited to a specific anatomy, but a general framework for CT reconstruction. Through Patient-Specific Optimization (PSO) on the patient’s own pCT, the model learns detailed anatomical priors and enables adaptation to temporal anatomical changes. Moreover, region-specific loss can further boost performance. Its robustness has been validated on complex head and neck data.

Domain Gap Problem [R1-3] We clarify that the Alignment Network (AN, Fig. 2) is designed for X-rays, so CT pretrained weights are not directly applied. We agree that a domain gap does exist. However, a shared latent semantic representation has been validated for reconstruction [1] since both rely on kV X-rays and similar attenuation. We thank Reviewer 1 for pointing out a promising future research direction regarding the domain gap.

Ill-posed Nature [R1-4] For the ill-posed nature, we emphasize that exact reconstruction is unattainable; our goal is to approximate the ground truth by leveraging priors, e.g., patient-specific priors from pCT. Experiment results showed our method achieved best performance by exploiting these priors.

Why not Registration or Limited-angle CT [R1-5, R3-4] Adaptive Radiotherapy (ART) requires updated 3D tissue density for dose optimization. 2D-3D registration is insufficient as it estimates only position. Limited-angle CT takes 10–30s to acquire 30–100 projections, causing image degradation from organ motion (e.g., peristalsis) and higher radiation.

Clinical Application [R1-6, R3-1/2] Radiotherapy is a cancer treatment using radiation to destroy tumors, delivered over multiple fractions across several weeks. In the standard ART workflow, a planning CT (pCT) is first acquired, on which the clinical target and radiation dose are determined. At each fraction, CBCT is acquired from the onboard kV system for patient setup and plan optimization (e.g., dose optimization) before treatment, which is time-consuming, high-dose, and increases complication risks. Therefore, our application directly utilizes the bi-planar X-ray mode of the same kV system to reconstruct CT volumes, offering a faster, lower-dose substitute for CBCT. Moreover, as the anatomy constantly changes over the treatment weeks, it is important to reconstruct CT at each fraction rather than relying on the pCT.

About the Evaluation [R1-7, R2-2/4] NeRF-based comparison. We clarify that INRR3CT is NeRF-based, which we quantitatively outperform quantitatively (e.g. +2.1 dB PSNR) and qualitatively. DRRs as synthetic X-rays. While we agree that real X-ray data is ideal, collecting such data is ethically prohibitive due to additional radiation solely for research. Therefore, evaluation with DRRs remains standard in this field [1,2]. Quantitative results and cohort sizes in Sec.3.4.This experiment used data from all 92 patients, and quantitative metrics align with qualitative findings, which surpass the runner-up method by 6.2% in HD95 and 2.3% in IoU.

About Hyperparameters [R1-8, R2-1] Weighting parameters. Loss weighting was set to balance the scales: λ₃=100, λ₄=0.1, and others=1.Unclear variables in Fig.2.AN in Fig. 2 uses M=4 encoder stages and N=3 decoder stages.

Difference with Previous LDM [R2-3] Unlike previous work, our training uses EA to suppress 2D ambiguity and PSO to exploit patient-specific priors. Architecturally, we propose AN for cross-reference information from orthogonal views.

All the above discussions will be added to the revision.

[1] Shen, L., Zhao, W., Xing, L.: Patient-specific reconstruction of volumetric computed tomography images from a single projection view via deep learning. Nat. Biomed. Eng 3, 880–888 (2019) [2] Ying, X., Guo, H., Ma, K., Wu, J., Weng, Z., Zheng, Y.: X2CT-GAN: Reconstructing CT from biplanar x-rays with generative adversarial networks. In: Conf. Comput. Vis. Pattern Recognit. pp. 10619–10628.(2019)




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    This paper proposes EA-LDM, an explicit-alignment latent diffusion model for patient-specific CT reconstruction from bi-planar X-rays. The reviewers recognized the potential clinical relevance of reconstructing or aligning patient-specific 3D anatomy from sparse intraoperative X-ray views, as well as the technical contributions of explicit alignment networks and patient-specific optimization. Some reviewers also considered the paper well structured, with useful ablation studies and competitive quantitative results against recent reconstruction baselines.

    However, the reviews also identified several important concerns. The clinical motivation and problem formulation are not fully clear, especially why full 3D reconstruction from two X-rays is preferable to conventional 2D–3D registration when preoperative CT is available. Reviewers also raised concerns about the ill-posed nature of reconstructing 3D CT from only AP/lateral X-rays, the lack of validation on real biplanar X-ray images, and unclear adaptation to specific anatomical regions. The use of pretrained CT-based latent representations for X-ray inputs also requires stronger justification given the domain gap. There are additional reproducibility concerns, including undefined hyperparameters, unclear variables in figures, and missing details about cohort sizes and loss weighting.

    Given the relatively positive scores from two reviewers and the overall quality of the submission, I recommend giving the authors a rebuttal opportunity. The authors should carefully address the reviewers’ concerns during rebuttal.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    Despite highly polarized post-rebuttal feedback (Accept, Accept, Reject), the authors’ rebuttal successfully dismantled the negative consensus by resolving core clinical and technical misunderstandings. First, they justified the reconstruction task over traditional 2D-3D registration by explaining that active radiation dose optimization strictly requires updated 3D tissue density volumes rather than mere spatial alignment. Second, they established true clinical utility, showing that pairing instantaneous bi-planar X-rays with historical pCTs accurately captures ongoing temporal anatomical variations (e.g., tumor regression) while entirely bypassing the high radiation dose and minute-long gantry rotations mandated by onboard CBCT or limited-angle scans. Third, they clarified that multi-view rendering baselines were not ignored, as the heavily outperformed baseline INRR3CT is itself a NeRF-based architecture. Backed by a 92-patient clinical cohort demonstrating superior numerical (PSNR/SSIM), visual, and downstream segmentation fidelity, this well-grounded framework satisfies publication standards and is recommended for acceptance.



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    Authors have adequately addressed comments.



Meta-review #3

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Reject

  • Please justify your recommendation.

    The clinical scenario targeted by this work remains somewhat controversial. In addition, many technical details and design choices are not clearly explained, and the overall organization of the manuscript needs substantial improvement. Given the considerable amount of revision required, I do not believe the paper is suitable for publication in its current form.



back to top