Abstract

Simultaneous reconstruction of sound speed and acoustic impedance in Ultrasound Computed Tomography (USCT) is essential for quantitative tissue characterization but remains a severely ill-posed inverse problem. Conventional Full Waveform Inversion (FWI) methods often suffer from cycle-skipping and parameter crosstalk due to non-convex optimization landscapes and limited low-frequency data. We propose InSPIRE, a framework for robust, unsupervised multiparameter inversion via sequential position-based implicit representations. Unlike discrete pixel-based approaches, we parameterize physical properties as continuous spatial functions using implicit neural representations. We construct a geometry-aware positional encoding with radially modulated Fourier features, imposing spatially adaptive spectral bias to resolve high-frequency details in the central imaging region where ring-array illumination is restricted. To disentangle parameter crosstalk, we introduce a dual-branch architecture with frequency-adaptive coordinate attention that dynamically adjusts feature importance across frequency stages, enabling independent reconstruction of sound speed and acoustic impedance. We devise a frequency-progressive sequential optimization strategy stabilizing convergence by updating the reference state with implicit noise regularization during bandwidth transitions. Extensive validation on synthetic phantoms and in vivo breast data demonstrates that InSPIRE outperforms conventional multi-scale FWI, yielding high-fidelity parametric maps with minimized crosstalk, superior convergence, and enhanced reconstruction quality.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/1332_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

N/A

Link to the Dataset(s)

N/A

BibTex

@InProceedings{ZenXia_InSPIRE_MICCAI2026,
        author = { Zeng, Xiaolu AND Yan, Weicheng AND Liu, Zhaohui AND Wu, Yun AND Tan, Hongrui AND Qiu, Wu AND Yuchi, Ming},
        title = { { InSPIRE: Multiparameter Inversion for Ultrasound Computed Tomography via Sequential Position-Based Implicit Representation } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16888},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    This paper presents an unsupervised INR-based framework for multiparameter ultrasound computed tomography (USCT), aiming to jointly reconstruct sound speed and acoustic impedance. The method combines a geometry-aware positional encoding with radially modulated Fourier features, a dual-branch architecture with frequency-adaptive coordinate attention, and a sequential frequency-progressive optimization strategy. Experiments on synthetic and in vivo breast USCT data show improvements over several FWI- and DIP-based baselines.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1.Joint reconstruction of sound speed and acoustic impedance is an important extension of quantitative USCT, since the two parameters carry complementary information about tissue properties and interfaces. The paper is also well motivated by two well-known challenges in waveform inversion, namely cycle-skipping and parameter crosstalk. 2.The geometry-aware encoding is well aligned with the ring-array acquisition setting, and the dual-branch design is a reasonable architectural choice for separating parameters with different spatial characteristics. Likewise, the frequency-progressive strategy is consistent with standard practice in stabilizing FWI optimization.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1.The current submission appears highly continuous with [1] at the levels of problem formulation, methodological perspective, and empirical presentation. Both works formulate USCT reconstruction as a subject-specific unsupervised INR/FWI problem, both exploit acquisition geometry in the coordinate representation, and both follow a very similar experimental narrative including synthetic studies, in vivo visualization, ablation analysis, and convergence behavior. While the present paper extends the target from single-parameter sound-speed reconstruction to joint sound speed and acoustic impedance estimation, the manuscript does not yet make sufficiently precise which aspects constitute a substantive methodological advance beyond this prior line of work. 2.Several components are introduced simultaneously: geometry-aware radial modulation, a dual-branch formulation, frequency-adaptive coordinate attention, and sequential optimization. However, the manuscript does not fully clarify which of these should be viewed as the primary scientific contribution, and which are better understood as design choices that support the move to multiparameter inversion. This makes the contribution boundary somewhat diffuse. 3.Since no in vivo ground truth is available, the validation mainly relies on CNR and qualitative structural consistency. This is understandable in the application setting, but it also means that the practical significance of the impedance estimates is somewhat less firmly established than the framing of the paper may suggest. [1] Wang Z, Yan W, Liu Z, et al. P²INR-FWI: An Implicit Neural Representation Method for Speed of Sound Image Reconstruction in Ultrasound Computed Tomography. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer Nature Switzerland, 2025: 420–430.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    1.Please clarify more explicitly how this work should be distinguished from Wang et al. [1] in methodological terms. In particular, which parts of the present framework are best viewed as conceptual advances, rather than extensions of an existing INR-based USCT reconstruction paradigm? 2.Please clarify what the paper considers to be its principal contribution: the move to multiparameter inversion itself, the dual-branch/FACA design, the geometry-aware radial modulation, the sequential optimization strategy, or the integration of these components. 3.Please discuss more directly the intended methodological relationship between this paper and prior INR-for-USCT reconstruction work, especially in terms of problem scope, technical inheritance, and claimed novelty. [1] Wang Z, Yan W, Liu Z, et al. P²INR-FWI: An Implicit Neural Representation Method for Speed of Sound Image Reconstruction in Ultrasound Computed Tomography. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. Cham: Springer Nature Switzerland, 2025: 420–430.

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    This is a relevant and technically competent paper on an important USCT reconstruction problem. The method is reasonably designed and the experimental section is broadly complete. My main reservation is not about technical quality per se, but about novelty positioning: the current version does not yet distinguish itself sharply enough from the most closely related prior INR-based USCT work, particularly Wang et al. [1].

  • Reviewer confidence

    Somewhat confident (2)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #2

  • Please describe the contribution of the paper

    The paper proposes InSPIRE, an unsupervised framework for multiparameter inversion in Ultrasound Computed Tomography (USCT). It introduces three main components: a geometry-aware positional encoding with radially modulated Fourier features to address ring-array illumination constraints; a dual-branch architecture with Frequency-Adaptive Coordinate Attention (FACA) to disentangle sound speed and acoustic impedance crosstalk; and a sequential frequency-progressive optimization strategy with noise regularization to prevent cycle-skipping.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    6.1 Novel formulation. The adaptation of Implicit Neural Representations (INRs) for USCT is mathematically elegant, specifically the radially modulated Fourier features that cleverly incorporate the physical constraints of ring-array geometry. 6.2 Effective Disentanglement. The proposed dual-branch FACA module successfully mitigates the velocity-density crosstalk, which is a notorious challenge in traditional Full-Waveform Inversion (FWI). 6.3 Clinical Feasibility. The unsupervised nature of the framework means it can operate without paired training data or ground truth labels. It demonstrates superior performance over standard FWI and DIP methods on both synthetic and in vivo data.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    7.1 On page 2, the manuscript states that recent work adapted “DIP” to USCT sound-speed inversion [22,26], but this conflates two rather different ideas: the original Deep Image Prior of Ulyanov et al. [22], which is a general untrained convolutional image prior for image restoration, and Yan et al. [26], which describes a USCT-specific untrained neural network integrated into an FWI framework for breast sound-speed imaging. The experimental section then refers to “DIP [26]” as a baseline, which makes it unclear whether the comparison is against the original DIP, Yan et al.’s USCT-specific UNN/FWI method, or an in-house variant. This ambiguity matters because baseline identity directly affects the interpretation and fairness of the reported gains. In addition, Wu et al. [25] is cited in the introduction, but the manuscript does not explicitly discuss it as the closest prior on multiparameter USCT FWI, nor explain why it is not included as an experimental baseline. I therefore suggest adding a short paragraph that explicitly distinguishes InSPIRE from (i) Yan et al., which focuses primarily on sound-speed-only untrained inversion within an FWI-style framework, and (ii) Wu et al., which addresses classical multiparameter USCT FWI without the proposed coordinate-based implicit representation, geometry-aware encoding, or dual-branch decoding. If Wu et al. is not used as a baseline, the reason should also be stated clearly. 7.2 A central claim of the paper is reduced multiparameter crosstalk. I think this concern should be evaluated more directly than by image quality metrics alone. A feasible addition would be a controlled synthetic experiment with spatially separated sound-speed and impedance anomalies, together with a simple leakage/contamination metric that quantifies off-target energy mapped into the wrong parameter image. A more advanced sensitivity/Hessian-style analysis would also be valuable, but I would view that as optional rather than mandatory for this paper. 7.3 The manuscript explicitly motivates the sequential frequency-progressive strategy as a way to stabilize convergence in the presence of cycle-skipping, and the current results do suggest improved convergence behavior. However, the evidence is still obtained under a relatively mild synthetic setup: homogeneous initialization, fixed 40 dB noise, and a single frequency schedule starting from 400 kHz. I therefore think the cycle-skipping claim would be much more convincing with one targeted stress test, preferably by varying the initial model error and/or the available low-frequency content. For example, the authors could perturb the homogeneous background model more aggressively, or repeat the experiment with the lowest frequency band removed. An SNR sweep would also be useful as an additional robustness check, but I would view that as secondary. 7.4 An important methodological clarification concerns the chosen parameterization. Eq. (1) is written in terms of sound speed and density, whereas the inversion variables are defined as sound speed and acoustic impedance. This is a valid choice in principle, since acoustic FWI can indeed be formulated in a velocity–impedance parameterization, but the manuscript should state this explicitly and explain how the forward model and gradients are implemented under (c,z). At present, it is unclear whether Eq. (1) is re-expressed directly in (c,z), or whether density is derived internally from z/cfor the solver. Because parameterization strongly affects sensitivity, trade-off, and crosstalk in multiparameter FWI, a brief justification for choosing (c,z) in this ring-array USCT setting would make the method section substantially clearer [5,11–13]. 7.5 The ablation study in Table 3 is currently difficult to interpret because its description appears inconsistent with the reported numbers. The text states that “the baseline uses a DIP network without Geometry-Aware Positional Encoding (GAPE) or attention modules (FACA),” which suggests that the ablation should start from a DIP-based configuration and then incrementally add the proposed components. However, the first row of Table 3 (“Base ×, GAPE ×, FACA ×, Freq-Progressive ×”) reports exactly the same quantitative results as the FWI baseline in Table 1, whereas the second row (“Base ✓, GAPE ×, FACA ×, Freq-Progressive ×”) matches the DIP results in Table 1.This creates ambiguity about what “Base” actually denotes and how the ablation sequence should be interpreted. I therefore suggest that the authors clarify whether Table 3 is intended to represent an incremental progression from FWI → DIP-style base network → +GAPE → +FACA → +frequency-progressive optimization, or instead an ablation entirely within the neural-prior family. As currently written, the table mixes baseline definitions in a way that makes the contribution of each component harder to assess. Since the ablation study is one of the main pieces of evidence supporting the paper’s claims, this inconsistency should be resolved explicitly. 7.6 The in vivo results are promising, but some of the manuscript’s clinical wording appears stronger than the current evidence supports. In particular, the conclusion states that the method provides information for “clinical diagnosis” and is “particularly suitable for clinical applications,” while the in vivo validation consists of only three volunteer cases evaluated mainly by ROI/background CNR and qualitative structural correspondence. I therefore suggest framing the in vivo study more explicitly as preliminary feasibility evidence, and softening claims of clinical diagnosis, clinical robustness, or clinical utility unless additional lesion-level, pathological, or reader-based validation is provided. 7.7 The manuscript provides enough information to understand the overall framework, but a few implementation details still seem necessary for reproducibility. In particular, it would be helpful to state explicitly either in the paper, the supplement, or released code: (i) the exact frequency-stage schedule used by InSPIRE itself, (ii) the noise scale σ in Eq. (10), (iii) the key Fourier-feature settings (at least the dimensionality and scaling of B), and (iv) the numerical specification of the shared encoder / dual-decoder architecture, or a clear pointer to an exact standard implementation if one is used. These items appear important because the reported gains depend on the interaction between architectural bias and staged optimization. 7.8 The geometry-aware radial modulation is intuitively motivated, but the current justification remains largely heuristic. Because the manuscript explicitly attributes the radial weighting to reduced illumination / spectral coverage in the central region of the ring-array setup, I would encourage the authors to provide at least one more direct supporting analysis or experiment. This does not necessarily require a heavy Hessian-based study; even a simple geometry- or radius-dependent sensitivity/illumination analysis, or an empirical radial ablation showing that central emphasis is specifically beneficial in this acquisition setup, would make the physical motivation more convincing.

    [1] Tancik, M., Srinivasan, P., Mildenhall, B., Fridovich-Keil, S., Raghavan, N., Singhal, U., Ramamoorthi, R., Barron, J., Ng, R. Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains. NeurIPS, 2020.[2] Shabtay, N., Schwartz, E., Giryes, R. PIP: Positional-encoding Image Prior. arXiv, 2022.[3] Hou, Q., Zhou, D., Feng, J. Coordinate Attention for Efficient Mobile Network Design. CVPR, 2021.[4] Bunks, C., Saleck, F. M., Zaleski, S., Chavent, G. Multiscale Seismic Waveform Inversion. Geophysics, 1995.[5] Operto, S., Gholami, Y., Prieux, V., Ribodetti, A., Brossier, R., Métivier, L., Virieux, J. A Guided Tour of Multiparameter Full-Waveform Inversion with Multicomponent Data: From Theory to Practice. The Leading Edge, 2013.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    I rated this paper as Weak Accept because it addresses an important and challenging problem in USCT and presents a reasonably well-motivated solution, but the current empirical and methodological support is not yet strong enough for a firmer recommendation. The main strength is the focus on joint reconstruction of sound speed and acoustic impedance, which is more clinically meaningful and technically more difficult than sound-speed-only imaging. I also found the overall framework coherent: the geometry-aware positional encoding is tied to the ring-array setting, the dual-branch decoder with FACA is motivated by reducing multiparameter crosstalk, and the frequency-progressive optimization is designed to improve convergence and mitigate cycle-skipping. The unsupervised, instance-specific formulation is also practically relevant for USCT, where paired ground truth is scarce. Within the reported experiments, the results are promising, with improvements over FWI, mFWI, and the neural-prior baseline on synthetic data, as well as encouraging in vivo results on three breast cases. My score remains borderline because several issues need clarification. First, the baseline definition is somewhat ambiguous, especially the use of “DIP” versus the USCT-specific untrained method in Ref. [26], which affects the fairness and interpretation of the comparison. Second, the claims regarding reduced crosstalk and improved resistance to cycle-skipping are plausible, but the current evidence is still indirect and would benefit from more targeted controlled experiments. Third, the paper should more clearly explain the chosen parameterization, since Eq. (1) is written in terms of sound speed and density whereas the inversion variables are sound speed and acoustic impedance. Fourth, the ablation study is difficult to interpret because the table and its description are not fully consistent. Finally, the clinical wording is somewhat stronger than warranted by the current evidence, given that the in vivo study includes only three volunteer cases and reproducibility-related implementation details remain limited. Overall, I view the paper as slightly above threshold, but my support would depend on the rebuttal adequately addressing these points.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #3

  • Please describe the contribution of the paper

    The paper proposes an unsupervised learning framework for ultrasound tomography image reconstruction using implicit neural representations with geometry aware position encoding.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    The paper introduces a lot of techniques and achieves the state-of-the-art imaging performance.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    The metrics on computation complexity such as parameter number and the reconstruction time are not compared. The authors emphasize the importance of Geometry-aware position encoding, yet the performance improvement by this is tiny as shown in ablation experiments.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not provide sufficient information for reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The proposed methods achieves superior performance in imaging. Yet comparison on computational complexity is not sufficient. Some details are missing for reproduction.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Author Feedback

We are grateful for the acceptance of our paper and the reviewers’ valuable feedback. We address the main concerns as follows: 1.Methodological Novelty (R1 and R2): We clarify how InSPIRE differs from Wang et al. and Yan et al. Wang’s P2INR-FWI uses point-wise MLP processing individual coordinates, while InSPIRE employs field-wise U-Net processing the entire coordinate field. This is critical for multiparameter inversion: our shared encoder captures global spatial context, enabling dual-branch decoders to independently reconstruct sound speed and impedance while mitigating crosstalk. Point-wise processing lacks spatial context to model distinct parameter characteristics. Yan et al. uses DIP mapping noise to images, whereas InSPIRE maps coordinates to physical parameters, providing continuous representation and stronger high-frequency modeling via Fourier features. Wu et al. [25] uses discrete FWI with optimal transport, a different paradigm from our coordinate-based continuous representation, making direct comparison less meaningful. 2.Experimental Design and Ablation (R2 and R3): We clarify Table 3’s ablation design. Row 1 is classical FWI. Row 2 uses DIP [26] as baseline for comparison with prior untrained neural USCT work. Rows 3-5 incrementally add our components. The transition from Row 2 to Row 3 includes shifting to coordinate-based input and adding GAPE. Coordinate-based formulation is a core design enabling our geometry-aware and frequency-adaptive modules. We note that coordinate-input U-Net without GAPE/FACA/frequency-progressive performs comparably to DIP, confirming coordinate-input provides similar regularization while enabling our proposed modules. This will be clarified in the final version. Regarding R3’s concern about GAPE’s modest gain, in synthetic experiments the ring-array geometry degradation is relatively mild, making radial weighting less pronounced. GAPE’s value is demonstrated through cumulative effects and becomes more critical in real scenarios with stronger illumination inhomogeneity. 3.Technical Details and Reproducibility (R2 and R3): Regarding the parameterization in c and z, our forward solver internally derives ρ=z/c, then solves Eq.(1) with gradients backpropagating to c and z. We chose this parameterization because impedance directly corresponds to tissue interfaces, facilitating dual-branch disentanglement. We will clarify this and provide complete implementation details in the final version. For computational cost, InSPIRE requires approximately 119 minutes on synthetic data using A100 GPU versus 170 minutes for DIP and 100 minutes for FWI. InSPIRE is faster than DIP and only moderately slower than FWI while delivering substantially improved reconstruction quality. 4.Clinical Validation (R1 and R2): We agree our in vivo study represents preliminary feasibility evidence on healthy volunteers rather than definitive clinical validation. Regarding impedance practical significance, acoustic impedance reveals tissue interfaces complementary to sound speed. While ground truth is unavailable, we validated structural consistency by comparing impedance gradients against USCT reflectivity maps in Fig. 5, demonstrating physically meaningful tissue boundaries. Additionally, quantitative characterization in healthy populations establishes baseline references for early screening protocols critical for preventive healthcare. InSPIRE shows promise for clinical applications, though further validation with patient cohorts and pathological correlation is needed.




Meta-Review

Meta-review #1

  • Your recommendation

    Provisional Accept

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    As all three reviewers have independently arrived at a positive recommendation, I extend my congratulations to the authors on the acceptance of their paper at this stage.



back to top