Abstract

Markerless intraoperative tracking of spinal anatomy via RGB-D sensors is a promising alternative to invasive bone-anchored fiducials, yet no existing clinical dataset provides independent subsurface ground truth for validating these methods. We address this gap with a clinically viable pipeline that derives per-vertebra 6-DoF pose from routine intraoperative C-arm fluoroscopy, co-registered to the RGB-D camera frame through a multimodal fiducial array detectable in both imaging modalities. Critically, the pipeline requires no invasive instrumentation contacting the patient, no additional radiation, or deviation from standard surgical workflows, relying solely on the fluoroscopic images already acquired during routine level checks. A calibrated transform chain links the vertebral pose recovered via multi-view 2D–3D CT-to-fluoroscopy registration to the RGB-D coordinate frame, enabling direct comparison of any markerless registration estimate against a subsurface reference. Ex vivo validation against optically tracked ovine vertebrae confirms end-to-end accuracy of 2.06 ± 0.29 mm and 3.35 ± 0.93°, sufficient to serve as independent ground truth for methods with substantially larger errors. We demonstrate the pipeline’s utility by evaluating two markerless registration strategies on clinical data spanning 11 vertebrae across 1,500 frames from three surgical sequences: the first such evaluation against an independent subsurface reference in a clinical setting. The proposed framework provides reusable infrastructure for benchmarking emerging markerless tracking methods on real surgical data.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/0829_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

https://github.com/condog101/BMTVG

Link to the Dataset(s)

CAD files and spine meshes: https://github.com/condog101/BMTVG depth videos: https://huggingface.co/datasets/zcbecda/BMTVG

BibTex

@InProceedings{DalCon_Bridging_MICCAI2026,
        author = { Daly, Connor AND Russo, Salvatore AND Ekanayake, Jinendra AND Elson, Daniel S. AND Rodriguez y Baena, Ferdinando},
        title = { { Bridging the Markerless Tracking Validation Gap: Intraoperative X-Ray Workflow Integration } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16893},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    The main contribution is the proposed pipeline for establishing “ground-truth” vertebral poses from pre-op CT and co-registered intra-op X-Ray and RGB-D images. It also claims to share the acquired dataset for benchmarking and share the pipeline design for transparency and reproducible studies.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    • A careful combination of existing novel methods
    • A collection of data and a contribution to a dataset for benchmarking, which is not limited to the proposed subsurface standard. In my opinion, the collected dataset is also valuable for analytical landmark evaluation.
    • A first-study of markerless tracking comparison using “established” X-ray poses, which is interesting to me.
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
    • The paper could have a stronger motivation for the proposed subsurface standard. The key citation on that is for skull-based applications, which is weakly relevant to the targeted application. The authors might want to provide more motivation for why it is important to have a subsurface standard benchmark dataset.
    • The paper could go further to provide actual subsurface measurements when comparing the markerless tracking methods, which was the goal. Yet, the current comparison felt short by presenting the measurement in the established X-ray poses. It would be more complete to demonstrate how the data generated by the proposed pipeline produces a subsurface standard comparison.
    • It seems to contain many goals in an eight-page paper. It might be better to focus on one main key contribution. From what I read, the key contribution appears to be the comparison, but it was not presented or discussed in detail. The authors might want to find a better balance among the three goals: a pipeline to generate data, a dataset for benchmarking the markless tracking algorithm, and the results of the comparison. In its current form, it is hard to tell which one is the focus.
    • The paper could also use more precise descriptions or clarifying terms that the authors intended to mean. e. g. , 1.It claims no additional instrumentation, but to me, putting a marker with a cover next to the exposure site is an additional piece of instrumentation. 2.It seems interchanging subsurface reference and vertebral pose or poses. It is confusing to me which one is the actual reference. A single vertebral pose or an average error among all vertebral poses, or a subsurface error measure? 3.It assumes the reader can interpret the error measures. It might be a standard, but it is still better to clearly describe the formula or provide a clear reference to the error measurement. To me, rotation errors normally have three directions, but it is simplified to one here. Also, there should be one error per vertebra, but mostly only one number is presented. 4.In section 3.1, it claims to evaluate each individual component. I read the individual components as: intrinsic projection evaluation, co-registration evaluation, tracking evaluation, and 2D-3D registration evaluation, but only the tracking evaluation was reported in the paper. It will help to have a table that clearly lists the error levels of each component, along with the overall error, so the reader can see it at a glance.
  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (2) Reject — should be rejected, independent of rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    This paper presents a pipeline for establishing vertebral “ground-truth” poses from pre-op CT and co-registered intra-op X-ray and RGB-D images, and it also intends to contribute a dataset as well as the material to deploy the proposed pipeline to generate the data for benchmarking markerless tracking methods. This work is interesting because it combines several well-performed methods in a careful and practical way, and the dataset itself appears valuable not only for the proposed subsurface benchmark but also for landmark evaluation.

    However, this paper raises several concerns. The motivation for subsurface standard benchmarking is not fully convincing, especially because the main supporting citation appears to come from skull-based applications. A more relevant standard measurement citation in spine-related surgery could further improve the motivation. In addition, although the stated goal is to enable subsurface-standard comparison, the paper does not go far enough to demonstrate actual subsurface measurements when comparing the two ICP methods, making the contribution feel incomplete. The paper also appears to pursue many objectives within these eight pages. A focused contribution would enhance the paper’s overall articulation. Finally, several terms and evaluation details would benefit from clearer definitions, including what is meant by “no additional instrumentation” (does it imply no additional setup or just a one-time setup?), and how to compute the error measurement used in the paper. Overall, the paper presents an interesting and potentially useful contribution, but its motivation, focus, and experimental presentation would benefit from further clarification and development.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The rebuttal addressed most of the major concerns. Although I cannot see the revised manuscript, if it is implemented, I would rerate it as Weak Accept. Thanks!



Review #2

  • Please describe the contribution of the paper

    This paper addresses an important validation gap in markerless intraoperative spine tracking: the lack of an independent subsurface reference in real clinical data. The authors propose a clinically feasible workflow that uses routine intraoperative C-arm fluoroscopy, preoperative CT, and a multimodal fiducial array to recover per-vertebra 6-DoF pose in the RGB-D camera frame without additional radiation, invasive instrumentation, or deviation from standard workflow. The key idea is to use multi-view 2D–3D CT-to-fluoroscopy registration to estimate vertebral pose, then connect this pose to the RGB-D frame through a calibrated transform chain. Ex vivo experiments on ovine vertebrae show end-to-end accuracy of 2.06 ± 0.29 mm and 3.35 ± 0.93°, and the framework is then used to evaluate two markerless registration strategies on 11 vertebrae across 1,500 clinical frames. In my view, the main contribution is not a new registration algorithm, but a practical and reusable validation framework that enables quantitative benchmarking of markerless spine tracking on real surgical data against an independent subsurface reference.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    • The paper tackles an important bottleneck in markerless spine tracking: lack of an independent subsurface ground truth on clinical data. This is a meaningful problem for translation, not just an incremental benchmark extension.
    • The proposed pipeline is well aligned with clinical practice because it leverages routine C-arm acquisitions already performed for level checks, while avoiding additional radiation and invasive fiducials. That practical framing is one of the most compelling aspects of the work.
    • The ex vivo validation in Table 2 provides a quantified error floor, and the clinical comparison in Table 3 reveals that the rigid method is rotationally more stable than per-vertebra ICP. This makes the paper valuable even beyond the specific methods tested.
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
    • Most technical ingredients are established: ArUco/PnP-based pose recovery, fluoroscopy calibration, and multi-view 2D-3D registration are standard components, and the paper itself cites prior work such as Markelj et al. on 2D/3D registration, Liu et al. on radiograph-to-CBCT registration, and Navab et al. / Jain et al. on C-arm calibration and tracking. The novelty is mainly the integration into a clinical validation workflow.
    • The clinical study includes only three sequences, 11 vertebrae, and 1,500 frames. That is sufficient for a proof of concept, but not yet enough to establish robustness across anatomy, workflow variability, image quality, or institutions.
    • The paper would benefit from more detailed statistics and ablations, for example on fluoroscopic view count/angular separation, failure cases of marker detection or 2D-3D registration, and uncertainty propagation from the reference pipeline into the final benchmark results.
  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    I place this paper slightly above the acceptance threshold because it solves a genuinely important evaluation problem for markerless spine tracking: obtaining an independent subsurface reference on real intraoperative data. That contribution is clinically relevant and, to my knowledge, new in this specific form. The paper is also pragmatic: it builds around routine fluoroscopy rather than requiring invasive fiducials or workflow-disrupting hardware, and the ex vivo validation provides a reasonable quantitative basis for using the pipeline as benchmark infrastructure. The clinical comparison is small, but still informative and already yields a useful conclusion about the relative strengths of rigid versus per-vertebra ICP registration. My main reservations are that the algorithmic novelty is limited, since many components are established, and that the clinical evidence remains narrow in scale. Overall, I think the paper is a valuable infrastructure contribution, but not yet a strong methodological advance.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    My initial concerns were mainly about ambiguity in the validation target, the meaning of the “subsurface reference,” the degree to which the marker array constituted instrumentation, and whether the reported errors allowed a fair interpretation of clinical usefulness. The rebuttal satisfactorily resolves these points without relying on inappropriate new experiments. In particular, the authors clarify that the reference is the full 6-DoF pose of each CT-segmented vertebra in the X-ray/C-arm frame, recovered by multi-view 2D/3D fluoroscopy-to-CT registration, rather than a superficial surface measurement. They also clarify that the marker array is placed beside the exposure under a sterile cover, does not contact the patient, and avoids the workflow disruption of invasive alternatives. The explanation of rotational-error computation and the separation between marker-tracking residuals and full end-to-end 2D/3D pose errors makes the quantitative results more interpretable. The remaining weaknesses—limited clinical scale and the fact that several technical components are established—are real but not fatal. The paper’s main value lies in integrating known components into a clinically feasible validation workflow for markerless intraoperative spine tracking, with released data/code and a clearly motivated gap. I therefore believe the contribution is above the MICCAI acceptance threshold.



Review #3

  • Please describe the contribution of the paper

    Evaluating markerless or stereovision based tracking systems for intraoperative registration purposes is challenging to validate intraoperatively. The authors propose a novel framework to provide ground truth evaluation of RGB-D based registration approaches for open spine surgery by leveraging a 2D to 3D per-vertebrae registration based on intraoperative flouroscopic imaging to preoperative CT. They novelly introduce a multimodal reference array consisting of radiopaque ArUco markers visible by both RGB-D cameras and fluoroscopic imaging to establish a common reference frame allowing co-registration of preoperative imaging, fluoroscopic imaging, and RGB-D imaging and establishes a transformation chain to evaluate downstream registration methods from RGB-D approaches. The authors provide experimentation detailing accuracy of their ground truth compared to an optical tracking system and demonstrate the clinical relevance of their approach by evaluating two RGB-D registration methods against their ground truth pipeline.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    S1: RGB-D based registration methods show much promise to reduce radiation exposure of open spine surgery. However, quantifying the accuracy of these methods without intraoperative CT is challenging and can be inhibitive due to added radiation exposure and workflow disruption clinically. The authors have established an approach that makes use of fluoroscopy in a novel way as a ground truth evaluation scheme of 3D RGB-D to preoperative imaging registration via a well-designed multimodal frame that is highly novel. S2: The authors will release the CAD models of the reference frame together with the source code for evaluation. By demonstration of a straightforward 3‑D‑printing pipeline (including best sterilization practice) this approach will be highly valuable for researchers who wish to adopt the method in their own labs as a benchmark for radiation reducing open spine registration methods and has a potential of setting a standard approach.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    W1.The authors claim that the 2‑D‑to‑3‑D registration achieves submillimeter precision and use a standard software package. This claim is used as a basis for their ground‑truth pipeline. However, they do not provide a quantitative assessment of this component or discuss this further. If the end-to-end registration analysis of the second validation experiment addresses this concern, further clarity is needed. Further, the accuracy of their RGB-D camera system is not separated out. While the first experiment may address this concern, there is ambiguity of this experimental description in Table 1.It is not evident whether the reported values are differences between optical‑tracking positions and the RGB‑D/X‑ray estimated poses, or relative pose differences between two modalities. Additionally, validating the C‑arm’s tracking of the radiopaque markers does not directly verify the accuracy of the subsurface 3‑D ground truth, because the 2‑D‑to‑3‑D step still introduces error that is not reported.

    W2. Authors state when evaluating the two registration methods that “While both methods achieve mean translational errors below 5mm, neither yet meets the ≤ 2 mm and ≤ 5◦thresholds cited for spinal navigation [22], underscoring both the value of independent pose evaluation and the headroom remaining for emerging methods.” However, they claim their ground truth method is justifiable because “The end-to-end pipeline, validated ex vivo against fiducial-tracked vertebral pose, yields a combined TRE of 2.06±0.29 mm and rotational error of 3.35±0.93, approaching the ≤ 2 mm and ≤ 5 thresholds commonly cited for spinal navigation systems [22]. Critically, this level of precision is sufficient to serve as independent ground truth for evaluating markerless tracking methods, whose registration errors are substantially larger, particularly in rotation.” This evaluation method seems to be at the very limit of the acceptable navigation accuracy of spine surgery. Should the ground-truth reference be more accurate than that it is evaluating? How would this compare to CT where voxel spacing may be much lower as a ground truth? Further comment and discuss around this seeming discrepancy is needed.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The authors are tackling a challenging problem to establish an easily utilized ground truth for markerless registration accuracy that can easily be deployed in a clinical setting without disruption to current clinical workflows. They have very cleverly identified a novel method. This paper has a great potential to set a new research standard.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    A novel method contribution that cleverly integrates technologies. Potential to be a strong standard of evaluating RGB-D methods



Author Feedback

We thank the reviewers and AC. All three reviewers recognized the framework’s contribution and the value of the released dataset. We address the AC’s four mandated points first.

1.(AC; R1) What is the “subsurface reference”, and where are subsurface measurements actually reported? Both “Subsurface reference” and “vertebral pose” denote one quantity: the 6-DoF pose of each CT-segmented vertebra in the RGB-D camera frame. It pertains to the pose of the entire vertebra, which is recovered by multi-view 2D-3D fluoro-to-CT registration, not just surface elements visible to the RGB-D camera (i.e. dorsal surface). Tab. 3 is the subsurface comparison: each markerless RGB-D estimate of the same per-vertebra pose is differenced against this reference. We will rename Tab. 3 columns “Subsurface Pose Error” and clarify the caption.

2.(AC; R1) Is a covered marker placed beside the exposure not itself instrumentation? Contribution 1 specifies “no instrumentation within the surgical field”. The marker array sits on a boom arm beside the exposure under a sterile cover; it never contacts the patient, requires no incision or bone anchor, and is placed/removed outside the sterile field. The invasive alternative (Liebmann’s SpineDepth: K-wired posts plus intra-op CT) is not permitted in observational clinical work. We will reword Contribution 1 to remove the ambiguity.

3.(AC; R1) How is rotational error computed, and why only one number per vertebra? It is the SO(3) geodesic angle between reference and estimated poses (Huynh, JMIV 2009), the same scalar used by Liebmann et al. (MedIA 2024). Per-axis decomposition adds no information for the worst-case bound, however per-vertebra Tab. 3 values and per-axis decomposition will appear in an appendix, and the formula will be defined in Section 3. 4.(AC; R1; R3) Errors from each pipeline stage are not separately reported, and Tab. 1 is ambiguous. In Tab. 1 the marker array is moved between four poses; for each transition the change in 6-DoF pose is measured both by the imaging modality and by an NDI infrared optical tracker (our reference). The reported values are each modality’s tracking residual against this commercial gold-standard, not absolute pose differences between modalities. This isolates each modality’s marker-tracking accuracy: 0.70 mm (RGB-D), 1.19 mm (AP), 0.96 mm (oblique). Tab. 2 then measures the full chain end-to-end vs. fiducial-tracked vertebral pose (2.06 mm, 3.35°). The Tab. 1→Tab. 2 gap (~0.9 mm, ~1.5°) characterizes the additional error from the 2D-3D stage. Tab. 1 and Tab. 2 captions will be rewritten to make these points explicit.

5.(R3) Is the reference accurate enough given the navigation threshold, and how does it compare with CT voxel spacing? A reference need only be better than the systems it evaluates, not better than the clinical threshold itself. Our reference (2.06 mm, 3.35°) is smaller than the per-method errors (4.5-4.8 mm, 4.5-8.4°) in both dimensions, and the 3.84° gap between the two methods’ mean rotational errors (per-vertebra ICP 8.36°, rigid 4.52°; Tab. 3) further exceeds the 3.35° reference uncertainty, so the relative ranking (rigid is more rotationally stable) is robust to reference uncertainty. Sub-mm voxels do not translate to sub-mm registration accuracy in practice: CT-based pipelines (Liebmann et al.) report 1.5 ± 0.8 mm, finer than ours but with the workflow disruption we avoid.

6.(R2) The clinical scale is sufficient for proof of concept but limited. The pipeline and evaluation framework are designed to scale: these cases form part of a large-scale ongoing clinical study (Section 1), with the released CAD, code, and data enabling additional sites to broaden anatomy, workflow, and institutional coverage.

7.(R1) The motivation rests on a skull-based citation. We will replace [1] (Balachandran, skull) with spine-specific [22] (Miller, 2 mm/5 deg threshold) and [16] (Liebmann SpineDepth, CT-based reference).




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    The paper presents a benchmark for markerless spine tracking from CT to X-ray by providing data with an independent subsurface ground truth vertebral poses. It uses multi-view 2D/3D registration to estimate vertebral pose, which is calibrated with an RGB-D camera via a multi-modal fiducial array. Reviewers highlighted the importance of the benchmark, which required significant effort and innovation to create. The rebuttal should address issues raised regarding the measurements reported for subsurface registration, the question of additional instrumentation, error metrics, and the need for clear presentiation of results for the error levels of each component of the system.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The rebuttal promises to clarify the presentation of results, resolve ambiguities such as the covered marker placement, error calculations. All reviewers rate the paper as “Accept.”



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The rebuttal addressed most of the concerns raised. The authors should include the proposed changes in their final camera ready version



back to top