Abstract

As surgery increasingly integrates advanced imaging, algorithms, and robotics to automate complex tasks, human judgment of system correctness remains a vital safeguard for patient safety. A critical example is 2D/3D registration, where even small misalignments can lead to surgical errors. Current visualization strategies alone are insufficient for reliable misalignment detection, highlighting the need for algorithmic decision-support. Using an AI framework for 2D/3D registration quality assessment, augmented with explainable AI (XAI) mechanisms to clarify model predictions, we conducted a user study (N=18) systematically comparing decision-making across three conditions: Human-only, Human+AI, and Human+XAI. We evaluated both objective performance measures (accuracy, sensitivity, precision, specificity) and subjective factors (workload, trust, understanding). Collaborative paradigms (Human+AI and Human+XAI) showed improved sensitivity, precision, and specificity compared to Human-only. Participants experienced significantly lower workload in collaborative conditions relative to the Human-only condition. Moreover, participants reported greater understanding of AI predictions in the Human+XAI condition than in Human+AI, although no significant differences were observed between the two collaborative paradigms in perceived trust or workload. Human-AI collaboration can enhance 2D/3D registration quality assurance, with explainability mechanisms improving user understanding. Future work should refine XAI designs to translate improved understanding into optimized decision-making performance.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/6000_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

N/A

Link to the Dataset(s)

N/A

BibTex

@InProceedings{ChoSue_HumanAI_MICCAI2026,
        author = { Cho, Sue Min AND Taylor, Russell H. AND Unberath, Mathias},
        title = { { Human-AI Collaboration for 2D/3D Registration Quality Assurance } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16889},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    This paper investigates how clinicians make accept or reject judgments when supported by AI assistance and explainability mechanisms in an assisted surgical context. The principal contribution is a controlled, within-subject evaluation comparing three decision-making conditions, human-only, human-plus-AI, and human-plus-explainable-AI, using objective performance metrics, including accuracy, sensitivity, precision, and specificity, alongside subjective human-factors measures encompassing cognitive workload, trust, and outcome understanding. The study provides evidence that collaborative human-AI paradigms improve operator understanding and reduce cognitive load on safety-relevant metrics relative to human-only decision-making, and that explainability enhances users’ understanding of model outputs even where improvements in objective performance over AI-only assistance are limited or non-significant. The paper contributes actionable evidence for the design of human-centered assurance workflows in safety-critical clinical AI systems, clarifying when and how AI explanations translate into improved operator decisions and reduced cognitive burden.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1) Clinical and CAI relevance: The topic addresses a question of substantial importance for safety-critical quality assurance in image-guided surgery. The paper focuses on 2D/3D registration, a task in which incorrect algorithmic alignment can result in severe consequences, including critical deviations in tool placement or implant positioning within delicate anatomical structures such as the spine. The clinical motivation is clearly articulated, with the authors appropriately identifying that current visualization strategies alone are insufficient to support reliable misalignment detection by human operators.

    2) Evaluation design: The study employs a rigorous within-subjects controlled user study design (N=18) comparing three decision-making conditions: human-only, human with AI assistance, and human with explainable AI (XAI) assistance. Counterbalancing is appropriately applied to mitigate order effects, and the sequence of AI-assisted conditions is randomized to reduce systematic bias. The evaluation encompasses both objective performance metrics, including accuracy, sensitivity, specificity, and precision, and subjective human-factors measures, specifically the NASA Task Load Index (NASA-TLX) for workload, trust, and operator understanding. This comprehensive approach is well aligned with clinical translation criteria, which emphasize user interaction, usability, and adoption readiness.

    3) Statistical reporting: The statistical methodology is robust and clearly reported. Normality of paired differences is assessed using the Shapiro-Wilk test to determine the appropriate choice between parametric paired t-tests and the nonparametric Wilcoxon signed-rank test. Family-wise error rates are appropriately controlled through the application of Holm-Bonferroni multiple-comparison corrections. The explicit reporting of means with standard deviations, p-values, and effect sizes, including Cohen’s d and rank-biserial r, across summary tables, substantially enhances the interpretability and transparency of the reported findings.

    4) Practical insights for real-world CAI systems: Beyond aggregate performance metrics, the paper provides actionable insights into how AI assistance and the addition of explainability mechanisms influence operator understanding and cognitive workload. The findings demonstrate that human-AI collaboration meaningfully improves safety-critical outcomes, with notable gains in both sensitivity and specificity, reflecting improved ability to correctly identify successful registrations and to dismiss failed ones. Furthermore, while AI assistance broadly reduces operator workload, the specific incorporation of XAI yields a statistically significant improvement in operators’ understanding of model predictions, a distinction with important implications for clinical deployment.

    5) Openness and ethical standards: The study explicitly confirms that institutional review board (IRB) approval and informed consent were obtained prior to participant involvement. Reproducibility is further supported by the accompanying algorithmic supplementary submission.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1) Generalizability of participants and task setting: The user study enrolls 18 participants drawn exclusively from a STEM background who, while representative of personnel involved in safety assurance of computer-assisted intervention (CAI) systems, are not the only primary end users in a clinical workflow. The absence of practicing clinicians or surgeons limits the external validity of the findings, as these groups may interpret explainable AI (XAI) overlays differently from technically trained non-clinicians and may exhibit distinct baseline trust levels shaped by clinical experience and professional accountability. The study is additionally conducted on a browser-based platform rather than within an integrated operating room environment, which does not capture the time pressure, physical demands, and cognitive distractions characteristic of real intraoperative conditions. The authors should more precisely define their intended operator profile in the introduction, and the discussion should explicitly address how the reported workload and trust outcomes might differ if the study were replicated with clinicians under realistic intraoperative constraints. 2) Case distribution realism and its effect on trust: The evaluation draws on a curated, balanced case set comprising equal numbers of true positives, true negatives, false positives, and false negatives. While this design ensures that participants encounter each AI outcome type, it artificially inflates the prevalence of AI errors relative to a natural clinical distribution, in which the model would be expected to perform correctly for the substantial majority of cases. As the authors acknowledge, repeated exposure to AI failures may suppress user trust and systematically bias subjective ratings. The finding that trust did not differ significantly between the AI-assisted and XAI-assisted conditions may therefore reflect this artificially elevated error prevalence rather than a genuine equivalence in trust response. The paper would benefit from a dedicated discussion of how this distributional choice likely influenced the NASA Task Load Index (NASA-TLX) and trust outcomes, and how the subjective findings might be expected to shift under a more representative real-world case distribution.

    3) XAI impact interpretation and methodological novelty: The paper applies established XAI methods, specifically Grad-CAM heatmaps and confidence scores, rather than proposing a novel explainability formulation. While the use of off-the-shelf techniques is appropriate for an application-focused contribution, a more substantive concern is that the XAI implementation, despite yielding a statistically significant improvement in subjective operator understanding, does not produce corresponding gains in objective performance metric, including accuracy, precision, and sensitivity, nor does it meaningfully reduce cognitive workload relative to the AI-only condition. This dissociation between improved understanding and unchanged decision accuracy raises questions about whether the selected explanation design provides operators with the type of actionable information required to improve task performance in this specific decision context. The authors should explicitly acknowledge that the primary contribution of this work lies in the human-factors evaluation rather than in algorithmic innovation, and the discussion should directly address the Grad-CAM design choices in relation to the absence of objective performance gains. Concrete directions for explanation refinement, such as interactive or conversational interfaces better suited to supporting decision accuracy, would strengthen the paper’s forward-looking contribution.

    4) Statistical power and high variance: The study employs a sample of 18 participants and reports substantial standard deviations across several objective performance metrics. For example, sensitivity in the human-with-XAI condition carries a standard deviation of 0.239.As the authors acknowledge, this combination of high within-condition variability and modest sample size may render the study insufficiently powered to detect smaller yet clinically meaningful differences between the AI-assisted and XAI-assisted conditions. The absence of statistically significant objective improvement in the XAI condition cannot therefore be interpreted as definitive evidence that XAI confers no performance benefit; the null result may reflect a Type II error rather than a true absence of effect. The authors should frame this limitation explicitly, characterizing the non-significant between-condition comparisons as potentially power-limited and recommending adequately powered replication as a priority for future work.

    5) Absence of deployment-centric metrics: The paper evaluates human-AI collaboration in terms of decision accuracy and subjective human-factors outcomes but does not report system-level metrics relevant to clinical feasibility, including inference time, computational scalability, and interface rendering latency. For a 2D/3D registration quality assurance system to be viable in image-guided surgery, the AI must deliver its assessment within a timeframe that does not disrupt the intraoperative workflow; if the combined inference and Grad-CAM generation pipeline introduces unacceptable latency, the practical utility of the system is substantially diminished regardless of its accuracy. The authors should report the average computational time required to generate AI assessments and XAI overlays as used in the study, and should briefly discuss whether the observed latency is compatible with intraoperative deployment requirements.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    System integration and cost-effectiveness: The paper demonstrates integration at a prototype interface level, which is appropriate for a controlled evaluation; however, it does not address the requirements for full deployment within an operating room system. No discussion is provided regarding the computational cost or resource requirements of implementing the AI and XAI components in practice. Briefly acknowledging these clinical translation barriers, including infrastructure compatibility, hardware requirements, and workflow integration, would meaningfully strengthen the broader translational context of the work.

    Several directions are recommended for future work. First, interactive XAI interfaces warrant investigation: the present study employs a static interaction model, specifically, the passive viewing of Grad-CAM overlays, to enable controlled evaluation. A productive next step would be to develop and evaluate iterative or conversational interfaces that allow operators to dynamically query the model regarding specific anatomical regions or uncertain predictions, with the aim of bridging the gap between the improved subjective understanding demonstrated in this study and objective decision accuracy. Second, behavioral and physiological tracking should be incorporated in subsequent studies to deepen understanding of the human-AI decision-making process. Gaze tracking and related physiological measures would provide quantitative insight into how operators visually search for misalignments and how attentional patterns shift in the presence of AI explanations. Third, ecological validation of trust calibration is warranted: as the balanced stimulus set, comprising equal numbers of true positives, true negatives, false positives, and false negatives, likely suppressed user trust by over-representing AI failures relative to a natural clinical distribution, a follow-up study using the model’s real-world error distribution is needed to obtain a more accurate assessment of trust calibration and perceived reliability under representative operating conditions.

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (5) Accept — should be accepted, independent of rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    This assessment reflects a balanced evaluation of the paper’s rigorous experimental design and clinically relevant application, weighed against its limitations in ecological validity.

    The principal strengths supporting this recommendation are as follows.

    First, the work addresses a substantive unmet need in image-guided surgery: operator-centered safety assurance for 2D/3D registration, a task in which reliable human detection of algorithmic misalignment is critical.

    Second, the evaluation methodology is strong for a conference-level clinical translation paper, employing a controlled within-subjects user study enrolling 18 participants that appropriately assesses both objective performance metrics and subjective human-factors outcomes, including the NASA Task Load Index (NASA-TLX), trust, and operator understanding.

    Third, the findings demonstrate that human-AI collaboration improves safety-relevant outcomes, with statistically significant gains in sensitivity and specificity relative to the human-only baseline, and that AI assistance meaningfully reduces operator workload. A further finding of note is that the addition of explainable AI (XAI) mechanisms produces a statistically significant improvement in operator understanding of model predictions.

    The following limitations temper these strengths.

    The user study enrolls STEM-background participants rather than practicing clinicians, which constrains the generalizability of findings to how surgeons would engage with the system under real operating room conditions. The evaluation employs a curated, balanced case set, comprising equal numbers of true positives, true negatives, false positives, and false negatives, which inflates the prevalence of AI errors relative to a natural clinical distribution and may bias subjective trust ratings. While XAI improves subjective understanding, it does not yield statistically significant gains in objective decision-making performance or trust relative to the AI-only condition. Finally, the modest sample size and high within-condition variance limit statistical power to detect smaller yet potentially clinically meaningful differences between the AI-assisted and XAI-assisted conditions.

    Notwithstanding these limitations, which are acceptable for an application paper at the conference stage, the methodology and actionable insights into human-AI collaboration in a safety-critical domain constitute a solid and timely contribution to the MICCAI community.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #2

  • Please describe the contribution of the paper

    This MICCAI submission studies human–AI collaboration for quality assurance (QA) of 2D/3D registration in image-guided surgery. The core idea is that, because small registration errors can have major clinical consequences, humans must be able to reliably accept/reject registration results—but visualization alone is often insufficient. The paper evaluates whether an AI “second opinion” and explainability (XAI) can improve human decision-making.

    Methodologically, the main paper reports a within-subject user study (N=18) comparing three conditions: Human-only, Human+AI (AI accept/reject), and Human+XAI (AI decision + confidence + Grad-CAM heatmaps). Outcomes include objective decision metrics (accuracy, sensitivity, precision, specificity) and subjective measures (NASA-TLX workload; perceived trust/understanding/helpfulness).

    The supplementary material details the underlying AI QA model: early fusion of fluoroscopy + DRR, a CNN backbone with spatial cross-attention, success defined as mTRE < 2 mm, and XAI via Grad-CAM plus conformal prediction for calibrated uncertainty. It is evaluated with leave-one-specimen-out CV on a pelvic fluoroscopy dataset.

    This fits MICCAI well at the intersection of image-guided interventions, quality assurance, uncertainty/explainability, and human-centered medical AI.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    Clinically relevant problem framing: QA for 2D/3D registration is safety-critical, and the paper’s emphasis on preventing false acceptance is well-motivated. Human factors component is a solid MICCAI angle: Evaluating how humans use AI support (not only model accuracy) is valuable and relatively underrepresented in classic registration papers. Clear experimental comparison of interaction paradigms: The three-condition design (Human-only vs Human+AI vs Human+XAI) is straightforward and interpretable; the UI is described and illustrated. Objective + subjective evaluation: Combining sensitivity/specificity with NASA-TLX and perceived understanding/trust gives a more complete picture than accuracy alone. Supplement adds technical depth: The AI model (early fusion + spatial cross-attention) and the use of conformal prediction are sensible design choices for calibrated decision support. The LOOCV protocol across specimens is appropriate for this setting. Results are directionally consistent: AI assistance improves objective metrics and reduces workload; XAI improves understanding ratings.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    Technical depth & novelty The human-study contribution is novel for registration QA, but the XAI mechanism itself (Grad-CAM overlays + confidence) is relatively standard. The paper would benefit from a stronger argument for why these specific explanations are expected to improve decision quality in this task (beyond general interpretability claims).

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The problem is important and the human-centered evaluation is relevant to MICCAI, but the novelty is limited relative to prior QA/verification + human visualization/uncertainty work, and the study design confounds (notably condition ordering and limited isolation of what XAI specifically contributes vs “extra signal”) leave the central claims not fully nailed down. With a strong rebuttal, I could see it moving to “Accept”.

  • Reviewer confidence

    Somewhat confident (2)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #3

  • Please describe the contribution of the paper

    In this paper, the authors conduct a user study that compares the performance of a group of users in assessing a 2D/3D registration task. The end users are asked to make an assessment without AI assistance, with AI assistance, and with explainable AI in the form of Grad-CAM–based heatmaps that highlight the image regions where the AI model places higher attention. Explainable AI is shown to provide added value across several metrics presented in the paper.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    • Rigorous analysis and clear presentation of findings.
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
    • The use of heatmaps to explain AI results in medical settings has been considered in past user-studies (see, for example https://www.nature.com/articles/s41598-024-82501-9).
    • The Algorithm/method used in this study is described in a parallel submission to the same conference. The current paper (focused on the user study) does not offer sufficient standalone contributions and depends on the acceptance of the method paper.
    • Although the participants have a STEM background, it is unclear whether individuals with such backgrounds would be responsible for making decisions related to 2D/3D registration in a clinical setting. The authors should provide clearer details on the relevant clinical workflows and decision-makers involved in accepting or rejecting AI outputs, and consider recruiting user profiles that more closely resemble those responsible for such decisions in clinical practice.
  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission has provided an anonymized link to the source code, dataset, or any other dependencies.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (2) Reject — should be rejected, independent of rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
    • The algorithm/method evaluated in this work is described in a dependent paper that is also submitted to the same conference. The current paper (focused on the user study) does not offer sufficient standalone contributions and depends on the acceptance of the method paper.
    • This work does not offer significant additional insights, either algorithmically or with respect to direct clinical practice.
    • The user study is not directly associated with a clinical workflow or practice.
  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Author Feedback

We thank all reviewers for their valuable feedback. We appreciate the recognition of our rigorous experimental design (R1), clinical relevance (R1, R2), and comprehensive evaluation combining objective and subjective measures (R1, R2). We feel encouraged by the acknowledgment that our work addresses a “substantive unmet need in image-guided surgery” (R1). Below, we address the points raised.

<Standalone Contribution and AI Model Details (R3)> We will revise the paper to be fully self-contained by incorporating essential details of the AI model architecture, including the early fusion approach, CNN backbone with spatial cross-attention, and Grad-CAM-based explainability mechanisms. The primary contribution remains the systematic evaluation of human-AI interaction paradigms for verification of 2D/3D registration, a distinct research question that require independent experimental design and human-factors analysis.

<Novelty of XAI Mechanisms (R1, R2, R3)> We agree that Grad-CAM and confidence scores are established techniques. The innovation derives from their systematic evaluation in human-AI collaboration for quality assurance of 2D/3D registration results. The finding that XAI improves subjective understanding without corresponding gains in objective performance offers valuable perspective on the gap between comprehension and actionable decision support. We will revise the introduction to highlight that the primary contribution is the human-factors evaluation, and expand discussion on XAI design directions (e.g., interactive interfaces).

<Participant Demographics (R1, R3)> We will update the introduction to more precisely define our intended operator profile. Our STEM-background participants represent engineers and specialists involved in CAI system validation, roles that are expanding as surgical systems become more advanced. We will expand the discussion to address how findings might differ with clinician participants under intraoperative constraints, identifying this as a priority for future validation.

<Balanced Case Distribution (R1)> We agree the balanced subset artificially inflates AI error prevalence. We will explicitly acknowledge how this likely influenced trust outcomes. We computed prevalence-weighted metrics to approximate real-world performance, and validating with actual error distributions remains important future work.

<Statistical Power (R1)> We acknowledge limited power from our sample size (N=18). However, we observed significant effects with substantial effect sizes (e.g., r=0.86 for understanding). This study identifies effect magnitudes enabling proper power analyses for subsequent studies. We will frame non-significant comparisons as potentially power-limited rather than definitive null results.

<Deployment Metrics (R1)> We appreciate the importance of deployment-centric metrics such as inference time and computational scalability. We will update the discussion to acknowledge these as important considerations and identify them as directions for future work as the system progresses toward clinical deployment.




Meta-Review

Meta-review #1

  • Your recommendation

    Provisional Accept

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    Given the overall positive reviews, I recommend acceptance. Please address the following concerns in the final version: the standalone contribution relative to the parallel method paper, the appropriateness of STEM participants versus clinical end users, the clinical workflow in which such decisions would be made, and the specific added value of XAI beyond standard AI assistance. The authors should also discuss the balanced case distribution, statistical power, and whether latency/deployment constraints are compatible with clinical use.



back to top