Abstract

Vision–language foundation models have shown strong potential in medical image analysis. Although foundation models for ultrasound imaging have recently emerged, the domain is particularly challenging due to severe speckle noise, acquisition variability, and subtle anatomical boundaries, leading to high inter-observer variability. Existing CLIP-based models rely primarily on global image–text alignment, limiting their sensitivity to clinically decisive local structures. We propose SonoCLIP, the first million-scale region-controllable fetal ultrasound vision-language foundation model that integrates segmentation masks as mask-channel visual prompts within an encoder, enabling joint global–local contrastive representation learning. To support scalable region–text alignment, we introduce a sigmoid-based pairwise contrastive loss that improves stability under large-scale supervision. We further curate a 1.44M-image multimodal fetal ultrasound dataset spanning 24 standard planes for large-scale pretraining. Extensive cross-center evaluations demonstrate that SonoCLIP achieves superior zero-shot transfer performance under both global and mask-guided inference, establishing a controllable and clinically oriented foundation model for fetal ultrasound analysis. Our code and data are available at: https://github.com/Harrison-one/SonoCLIP.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/5124_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

https://github.com/Harrison-one/SonoCLIP

Link to the Dataset(s)

Fetal24: https://drive.google.com/file/d/1mYcavqKTCb70yQVWYgPI5NIky6Cb-cFP/view Fetal6: https://drive.google.com/file/d/1ye2BXpeWvUuGIbykQ-bjt14HgNKg6qMU/view Fetal5: https://drive.google.com/file/d/1hJl7L79lYRXxXGC9D3KtKcGWt-O8gJuT/view

BibTex

@InProceedings{SuHan_SonoCLIP_MICCAI2026,
        author = { Su, Hang AND Sun, Chao AND Li, Zhaofan AND Hu, Wei AND Liu, Juhua AND Du, Bo},
        title = { { SonoCLIP: Mask-Guided Region-Aware Vision–Language Pretraining for Fetal Ultrasound Analysis } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16879},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    The paper proposes a fetal ultrasound foundation model designed to learn region-aware representations, with the goal of improving performance across downstream fetal ultrasound tasks. A central aspect of the method is the use of a mask-aware design and a sigmoid pairwise contrastive loss to better capture spatially relevant clinical information during representation learning.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    A key strength of the paper is that it focuses on fetal ultrasound, which is an important and clinically relevant imaging domain.

    The idea of learning region-aware representations for this modality is also interesting, particularly if such representations can support multiple downstream tasks within a unified framework.

    In addition, the paper includes experiments across downstream applications (classification and segmentation), which is valuable for assessing the practical utility of the proposed model. The reported improvement on segmentation is also potentially promising, especially if the method is able to transfer effectively across tasks.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    One point that would benefit from further clarification is the paper’s positioning relative to prior fetal ultrasound foundation models, especially FetalCLIP. Since FetalCLIP seems like a very natural work to compare this work to, it would be helpful for the paper to explain more explicitly in more depth how the proposed method differs conceptually and technically. Relatedly, this also makes it difficult to assess the claim that the method is the first foundation model for fetal ultrasound, and I think the paper would be strengthened by a more careful discussion of prior work (FetalCLIP) in this space.

    A second point is that some of the experimental findings would benefit from more intuitive explanation. For example, in Figs. 2 and 3, it is not entirely clear how the reader should interpret the comparison. Are all the models trained on fetal ultrasound data? It is perhaps not unexpected that a model trained on fetal ultrasound data would outperform models that were not trained for that domain. Additional discussion of what these figures demonstrate would help clarify the contribution.

    I also think the paper would benefit from more intuition around why the proposed design choices are effective. In particular, it would be useful to better explain why the sigmoid pairwise contrastive loss works well in this setting, why the mask channel provides such a benefit, and why the model performs particularly strongly on the downstream segmentation task. The empirical results are appreciated, but more interpretive explanation (discussion) would help the reader understand what aspects of the method are leading to the gains.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The empirical results are promising and relevant, but the paper would benefit from clearer positioning against FetalCLIP and more intuition for why the proposed design choices help.

  • Reviewer confidence

    Somewhat confident (2)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    I appreciate the authors’ rebuttal and thank them for the clarifications. I am comfortable recommending acceptance. My remaining suggestion is to revise the abstract’s claim that “no dedicated foundation models exists for fetal ultrasound”, as this appears too strong in light of FetalCLIP. I think the wording used in the rebuttal is more precise and would better position the contribution (“SonoCLIP is the first million-scale region-controllable fetal-US VLM”). With that adjustment, I am happy to support acceptance.



Review #2

  • Please describe the contribution of the paper

    This paper proposes SonoCLIP, a vision-language foundation model for fetal ultrasound, which uses a mask-channel visual pathway to introduce region-aware learning. In this method, global image-text alignment is combined with mask-guided region-text alignment, and the standard softmax contrastive objective is replaced with a sigmoid pairwise loss function. Moreover, the authors present a large multimodal fetal ultrasound dataset containing 1.44 million images, masks, plane labels, gestational age information, and structured captions. The experiments show significant improvements on cross-center zero-shot classification compared to CLIP, UniMed-CLIP, and FetalCLIP, as well as on smaller public datasets related to linear probes.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    Following are some of the key strengths of the paper.

    • In this paper, a clinically relevant and important problem is addressed. Fetal ultrasound is a challenging field because of image noise, operator dependence, and subtle anatomical structures. Therefore, a dedicated foundation model is indeed useful.

    • The method is simple and easy to follow as the idea of adding a mask as an extra visual prompt is intuitive, and the design choice of initializing the mask branch so that it does not affect the original CLIP behavior at the beginning is thoughtful and technically clean.

    • The dataset’s scale is one of the strongest features of the analysis. Pretraining on 1.44 million fetal ultrasound images with plane labels, masks, and text descriptions is a meaningful contribution by itself, especially in a specialized medical modality.

    • In ultrasound, mask-channel mechanisms are useful for focusing on local structures due to their simplicity and intuitive nature.

    • Compared with current baselines, the cross-center zero-shot benchmark shows substantial gains, especially on this task.

    • The paper is also generally well organized and easy to follow.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    Following are some of the weaknesses of the paper.

    • The mask-guided inference at test time is not clearly define and this is a critical point here that needs to be clarified. There is a very large gain reported when masks are used during inference, but it is not clear whether these masks are manually annotated, automatically predicted, or where these come from during the testing. . If they are ground-truth masks, then this setting is less practical and not directly comparable to standard zero-shot classification baselines.

    • Secondly, most evaluations report values without confidence intervals, repeated runs, or significance testing. Ablation studies are also limited, and they do not isolate the effects of mask quality, the impact of attention blocks, or rationale for using sigmoid based objective. These ablations are not the final but showing them would’ve better make the claim of SonoCLIP stronger in comparison to models like AlphaCLIP or SigLIP.

    • Furthermore, the main weakness appears to be the moderate novelty of the approach. As the mask-guided visual pathway is like promptable CLIP-style model (AlphaClip), and the sigmoid-based contrastive loss is also derived from SigLIP, the main contribution appears to be domain-specific integration and large-scale dataset curation rather than a fundamentally new method. Despite this limited novelty, the claims should be presented in a more detailed manner. Also, it would’ve been better to show comparison with AlphaCLIP as well as to SigLIP so as to show how the domain-specific integration helped in learning useful details.

    • Additionally, the paper lacks details about the dataset. In addition to the total number of images, the report does not detail the number of patients or studies, the distribution of scanner vendors, or the percentage of manually verified masks. Due to the high degree of correlation between images in ultrasound domain, assessing the scale and generalization solely from image count is not enough.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    Some key points are mentioned in the weakness section. Addressing them will make the work more stronger as currently I see it as an extension of AlphaCLIP-style method for ultrasound domain.

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper addresses an important problem, is well motivated, and shows strong results in a clinically meaningful ultrasound setting. The paper is especially convincing in showing that a ultrasound specific foundation model can transfer well across datasets and centers. The cross-center zero-shot results are a strong point but I have some concerns about method novelty, the unclear mask-based inference setup, and the lack of detail about data curation and annotation quality. So the overall overview of the paper is positive, but I would like to expect a strong rebuttal for the main weaknesses mentioned.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The authors addresses most of my major concerns. They clarify that mask-guided inference uses ground-truth masks and should be interpreted as a mask-assisted upper bound, while the fair zero-shot setting is the “w/o mask” version. This resolves the main ambiguity about test-time masks. They also provide stronger evidence through AlphaCLIP comparisons, confidence intervals, significance tests, and additional ablations for sigmoid loss and mask usage. The dataset details are improved with patient count, scanner information, and mask verification percentage, although more complete statistics would still be useful.

    My remaining concern is that the methodological novelty is moderate, since the approach builds on AlphaCLIP/SigLIP-style ideas. However, the fetal ultrasound-specific integration, region-level supervision, scale of the dataset, and improved empirical evidence make the contribution sufficiently valuable. Overall, the rebuttal substantially strengthens the paper, so I would revise my recommendation to Accept.



Review #3

  • Please describe the contribution of the paper

    This work presents a domain-specific foundation model for fetal ultrasound. It provides a comparison with the latest publicly available FetalCLIP, along with ablation studies. The proposed model demonstrates good performance on downstream classification and segmentation tasks, and its zero-shot classification results are particularly encouraging.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    The dataset used in this work is large in scale and is accompanied by corresponding segmentation masks. It contains 1.44 million fetal ultrasound images covering 24 standard anatomical planes. The construction of such a dataset clearly involves substantial effort and provides strong support for potential clinical applications in real-world settings.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    The proposed method still seems to follow existing paradigms. The authors attribute the performance gain of SonoCLIP (w/o mask) over FetalCLIP to the use of mixed global- and region-level supervision. It remains unclear whether the “region-level supervision” here specifically refers to mask-based supervision. This point should be clarified further, particularly to explain how such supervision contributes in the setting without mask guidance.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The work demonstrates clear clinical relevance. In particular, the authors’ plan to release a large-scale test set is likely to be valuable for the community and could help advance research in this area. That said, the contribution of the paper seems to rely more heavily on the scale and value of the dataset, while the methodological novelty appears relatively limited.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Author Feedback

We thank the reviewers for recognizing SonoCLIP’s clinical relevance and dataset value. We address the main concerns and will revise the manuscript.

1.Clarification of model comparison and novelty (R1, R2). We will clarify that our claim concerns the lack of formally published fetal ultrasound foundation models, while FetalCLIP is a relevant recent preprint/concurrent work. Unlike FetalCLIP, which mainly focuses on global image-caption alignment, SonoCLIP is pretrained on 1.44M-image multimodal fetal ultrasound dataset and introduces mask-channel prompts with mixed global/region-text supervision. To our knowledge, this makes SonoCLIP the first million-scale region-controllable fetal-US VLM. We will also add AlphaCLIP (without DWConv) ablations in Top-1/Top-5: AlphaCLIP w/o mask 47.09/88.76, w/ mask 68.44/97.34; SonoCLIP w/o mask 50.89/87.36, w/ mask 85.01/99.01. 2.Interpretation of comparison experiments in Figs. 2–3 (R1). CLIP is natural-image pretrained, UniMed-CLIP is general-medical, and FetalCLIP/SonoCLIP are fetal-US-specific. Thus, the key evidence is not fetal training beating generic models, but SonoCLIP improving over the closest fetal baseline. In Table 1, SonoCLIP w/o mask improves over FetalCLIP by +18.60/+11.22 Top-1/Top-5, showing that mask-based region supervision improves global inference. Mask-guided inference adds +26.63/+4.54, quantifying localization benefit. Fig. 3 evaluates cross-dataset transfer and dense prediction.

3.Training-time region supervision and test-time mask guidance (R2, R3). We will distinguish these settings clearly. “Region-level supervision” denotes mask-based region-text alignment during pretraining, while “mask guidance” refers to whether a mask is used at inference. “w/o mask” uses an all-one mask at inference, so no anatomical mask is provided and it is the fair zero-shot comparison to CLIP, UniMed-CLIP, and FetalCLIP. “w/ mask” uses ground-truth anatomical masks and will be reported as a mask-assisted upper-bound, not the main deployment protocol. Thus, the w/o-mask gain over FetalCLIP comes from anatomy-aware representations learned during mask-based pretraining, not test-time masks.

4.Confidence intervals, significance testing, and ablations (R2). We added bootstrap 95% CIs and paired McNemar tests. Table 1 Top-1/Top-5 CIs: CLIP [9.83,11.26]/[28.41,30.78], UniMed-CLIP [15.68,17.71]/[48.87,51.81], FetalCLIP [38.29,41.27]/[81.97,84.52], SonoCLIP w/o [57.06,59.69]/[93.69,95.23], SonoCLIP w/[83.97,86.02]/[98.65,99.33]. McNemar: SonoCLIP w/o vs. FetalCLIP p=4.46e-47; SonoCLIP w/ vs. w/o p=2.08e-158.Table 2 Acc CIs: CLIP [85.15,88.88], UniMed-CLIP [81.49,85.62], FetalCLIP [93.12,95.43], SonoCLIP w/o [95.23,97.30], SonoCLIP w/ [98.81,99.68]. McNemar: SonoCLIP w/ vs. FetalCLIP p=3.17e-12; SonoCLIP w/ vs. w/o p=3.24e-08.We will report these and clarify that SonoCLIP adds no extra attention block; isolated factors are mask channel, sigmoid loss, and fetal-specific region-text supervision.

5.Further clarification of sigmoid loss, mask channel, and segmentation gains (R1, R2). Our supervision mixes global and local mask captions. Sigmoid pairwise loss optimizes each image-text pair independently, fitting heterogeneous global/local positives better than batch-coupled softmax. Fig. 2 supports this: SigLoss improves Top-1/Top-5 from 50.89/87.36 to 58.38/94.47 w/o mask, and from 67.83/94.29 to 85.01/99.01 w/mask. The mask channel preserves anatomy- aware local cues, explaining stronger zero-shot matching and segmentation transfer.

6.Dataset statistics and annotation quality (R2). We will add patient-, vendor-, and annotation-level statistics to avoid relying solely on image count. The dataset currently includes 36,482 patients, mainly acquired from GE Voluson E8/E10 systems, and approximately 55% of masks have been manually verified or curated. More detailed dataset information will be disclosed in the revised version after data curation is completed.




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    Two reviewers lean weak accept and one weak reject. All three agree the work is clinically relevant, the dataset effort is substantial, and the empirical results, especially for cross-center zero-shot transfer, are promising.

    The authors need to clarify much more carefully how SonoCLIP differs from prior fetal ultrasound models, especially FetalCLIP, and from more general mask-guided CLIP-style approaches such as AlphaCLIP. Right now, the novelty appears to some reviewers to be more in the domain adaptation and dataset scale than in the method itself.

    The mask usage at inference needs to be stated very clearly. If masks are required at test time, where do they come from, and how practical is that setting? If ground-truth masks are used, then the comparison to standard zero-shot baselines becomes difficult to interpret.

    There are also requests for better ablations and dataset detail, including the role of the sigmoid objective, the contribution of the mask channel, the number of patients and studies rather than just images, scanner diversity, and annotation quality.

    These are central enough questions that rebuttal is appropriate before a final decision and it could go either way.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The paper proposes SonoCLIP, a mask-guided, region-aware vision-language model for fetal ultrasound. The work is clinically relevant and supported by a substantial dataset effort, with promising cross-center zero-shot transfer and downstream segmentation results. The main concerns were moderate methodological novelty, unclear relation to FetalCLIP and AlphaCLIP, and ambiguity about whether masks are required at inference.

    The authors clarify that the fair zero-shot comparison is the w/o-mask setting, while the w/-mask setting should be interpreted as a mask-assisted upper bound. They also clarify the distinction between training-time region supervision and test-time mask guidance, provide AlphaCLIP comparisons, confidence intervals, significance tests, and additional ablations for sigmoid loss and mask usage. Dataset details, including patient count, scanner type, and mask verification percentage, further strengthen the submission.



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    Reviewers are overall positive and consider the paper clinically relevant, with substantial dataset value and promising empirical results. The rebuttal clarified the role of masks at inference, explained that the mask-assisted setting should be interpreted as an upper bound rather than the main deployment setting, and strengthened the evaluation through clearer zero-shot interpretation, added ablations, significance tests, and improved dataset details.



Meta-review #3

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    I recommend acceptance. All three reviewers are positive after the rebuttal, with R2 moving from Weak Reject to Accept once the authors clarified that mask-guided inference uses ground-truth masks as an upper bound. At the same time, the fair zero-shot setting is the “w/o mask” version. The methodological novelty is moderate, as the approach adapts AlphaCLIP-style mask prompting and a SigLIP-style loss to fetal ultrasound, but the scale, consistent cross-center gains, and released resources make the contribution valuable. Two corrections are required in the camera-ready: first, the abstract’s claim that “no dedicated foundation model exists for fetal ultrasound” is inaccurate and contradicts the paper’s own citation of and comparison against FetalCLIP, so it should be revised to the more precise wording proposed in the rebuttal; second, the FetalCLIP segmentation result in Table 3 is lower than even natural-image CLIP, which is anomalous and should be explained or qualified to ensure the comparison is fair.



back to top