List of Papers Browse by Subject Areas Author List
Abstract
Accurate 3D anatomical segmentation is essential for clinical diagnosis and surgical planning. However, automated models frequently generate suboptimal shape predictions due to factors such as limited and imbalanced training data, inadequate labeling quality, and distribution shifts between training and deployment settings. A natural solution is to iteratively refine the predicted shape based on the radiologists’ verbal instructions. However, this is hindered by the scarcity of paired data that explicitly links erroneous shapes to corresponding corrective instructions. As an initial step toward addressing this limitation, we introduce CoWTalk, a benchmark comprising 3D arterial anatomies with controllable synthesized anatomical errors and their corresponding repairing instructions. Building on this benchmark, we further propose an iterative refinement model that represents 3D shapes as vector sets and interacts with textual instructions to progressively update the target shape. Experimental results demonstrate that our method achieves significant improvements over corrupted inputs and competitive baselines, highlighting the feasibility of language-driven clinician-in-the-loop refinement for 3D medical shapes modeling. Code and dataset will be released at https://github.com/HINTLab/CowTalk.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/0740_paper.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to the Code Repository
https://github.com/HINTLab/CowTalk
Link to the Dataset(s)
https://zenodo.org/records/20113566
BibTex
@InProceedings{XieKan_Refining_MICCAI2026,
author = { Xie, Kangxian AND Yang, Jiancheng AND Pinter, Nandor AND Wu, Chao AND Bozorgtabar, Behzad AND Gao, Mingchen},
title = { { Refining 3D Medical Segmentation with Verbal Instruction } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 16883},
month = {September},
page = {pending}
}
Reviews
Review #1
- Please describe the contribution of the paper
This paper introduces CoWTalk, a benchmark for language-guided 3D medical shape refinement, and proposes an iterative framework that refines segmentation outputs using natural language instructions. The work aims to enable human-in-the-loop correction of segmentation results.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
(1) Novel problem formulation: The idea of using natural language to iteratively refine 3D segmentation is original and aligns with emerging vision-language paradigms. (2) Dataset contribution: The CoWTalk benchmark provides paired data of corrupted shapes and corrective instructions, which is currently lacking in the medical domain. (3) Human-in-the-loop perspective: The framework reflects realistic clinical workflows where experts provide feedback. (4) Multi-modal modeling: The integration of shape representations and textual instructions is reasonable and consistent with current vision-language modeling practices.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
(1) The proposed dataset relies on synthetically generated segmentation errors and LLM-based instructions, which may not fully reflect real clinical error distributions and physician communication patterns. (2) The evaluation is restricted to vascular structures, and it remains unclear how well the method generalizes to other anatomical targets or more complex geometries. (3) The proposed model largely builds upon existing multi-modal components (e. g. , attention-based fusion and implicit decoding), and the level of architectural novelty is relatively limited. (4) The evaluation is mainly conducted on the same synthetic data distribution used for training, which may lead to overly optimistic results. The limited validation on real-world cases is not sufficient to demonstrate practical applicability.
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The paper explores a highly novel direction and introduces a valuable benchmark. However, the reliance on synthetic data and the absence of comprehensive real-world validation (beyond limited qualitative examples) weaken the evidence for its practical impact.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #2
- Please describe the contribution of the paper
The primary contribution of this paper is benchmark for instruction-following medical shape refinement. It comprises 3D arterial segmentations featuring controllable, synthesized anatomical errors, categorized into global types (such as thinning, thickening or missing segments) and local types (including local thinning, thickening, shortening, disconnection, fragmentation). Each entry includes a textual description of the errors alongside instructions of their improvement. This dataset is designed to serve as the benchmark of iterative, human-in-the loop corrections of medical segmentations guided by verbal instructions. Furthermore, the authors propose and evaluate an end-to-end vision-language iterative framework of their own design to address this task. The benchmark builds upon existing open-access dataset of Circle of Willis.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
1.Novel Benchmark for the Medical Domain: The paper introduces the first benchmark specifically designed for instruction-following medical shape refinement. The work addresses a significant gap as most instruction-guided editing research has been confined to general-purpose or non-medical objects. 2.Comprehensive Comparative Analysis: The evaluation compares the proposed framework against a wide range of baselines. These include CNN-based refinement methods (such as nnUnet v2) utilizing both shape-only and image plus shape inputs and general-purpose text-guided shape editing methods. Additionally, solution was evaluated on human (non-medical researcher and neuroradiologist) instructions to show its potential utility in real-world settings.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.Ambiguity Regarding Open Availability: The authors do not state whether the errors generation code, the benchmark dataset or the proposed method’s implementation will be made publicly available. Given that the primary contribution of this work is a new benchmark, its utility and impact on the field are dependent on the accessibility of these resources. 2.Issues with Figure Clarity: Several figures are difficult to interpret in their current form, in particular: Figure 2: the background choice makes the plotted points poorly visible, Figure 5: The text instructions are difficult to read. Figure 1: Human icon in the second row of the image is ambiguous.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not provide sufficient information for reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
My recommendation of Weak Accept is primarily driven by the potential impact and utility of the proposed benchmark for the research community. The recommendation is qualified as “Weak” due to the current lack of clarity regarding the public availability of the code and dataset, as well as several issues with figures clarity.
- Reviewer confidence
Somewhat confident (2)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #3
- Please describe the contribution of the paper
This manuscript provides a new 3D dataset, including paired medical forms and natural language geometric optimization descriptions. It also uses doctors’ written instructions as prior knowledge to correct the 3D vascular models that were incorrectly segmented by the other methods.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
1.A new and complete 3D dataset. 2.An intuitive and reasonable motivation: The automatic segmentation of AI will sometimes make obvious errors. At this time, the prior knowledge of doctors can promptly detect such errors, so medical textual instructions are used to correct them.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.The title mentions 3D medical segmentation, but the dataset only consists of Willis artery images. The title is somewhat exaggerated. 2.It does not clearly describe how the instructions are used in the reasoning process, whether iterations are required during the reasoning stage, and when the iterations should stop.
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not provide sufficient information for reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(5) Accept — should be accepted, independent of rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The paper has an intuitive motivation and a large‑scale dataset.
- Reviewer confidence
Somewhat confident (2)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Author Feedback
We thank the reviewer for the recommendation and comment. We make point-to-point response to each reviewer’s “weakness” comment. Reviewer 1: 1.We agree with the assessment. To mitigate this issue, we designed the error synthesis and instruction generation pipelines to maximize structural and linguistic diversity. Nevertheless, the generated data may still follow learnable distributions. We therefore view the current benchmark as an initial proof-of-concept for instruction-guided segmentation refinement and will clarify this limitation in the final paper. 2.We agree with the assessment. The current benchmark focuses on tubular vascular structures because centerlines provide a controllable representation for error synthesis and instruction generation. Extending to non-tubular anatomies is more challenging due to the lack of canonical spatial representations. In future work, we plan to leverage SOTA LLM/VLMs to model real-world 3D prediction errors and enable instruction generation for arbitrary anatomies. 3.We agree that the proposed architecture builds upon existing modules. However, extensive experimentation with alternative designs showed that many prior shape manipulation methods for general 3D objects fail on highly irregular medical shapes, as reflected in our results. We believe the proposed end-to-end framework provides a strong foundation for future work on complex medical shape manipulation.
Reviewer 2: 1.We acknowledge the importance of open-sourcing the dataset and generation pipeline. The complete dataset has been anonymously released at kaggle https://www.kaggle.com/datasets/daxianzzz/cowtalk, and we will include the github repo link for camera ready version, and release the full code and checkpoints before the conference. 2.Due to page limits, Figure 2 could not be substantially enlarged, but we increased its size slightly and added more explanatory captions for clarity. For Figure 5, some text appears small because of the amount of information shown; however, the PDF is lossless, so all details remain clear when zoomed in. We have also redrawn Figure 1 to significantly improve the clarity of the workflow illustration.
Reviewer 3: 1.We agree that the current study is evaluated specifically on CoW, and therefore may suggest a broader level of generality. At the current stage, we are unable to modify the title due to submission constraints. However, we will revise the manuscript to more explicitly acknowledge this limitation and clarify that the present work should be viewed as an initial proof-of-concept for language-guided refinement of 3D medical segmentation shapes. 2.We would like to clarify that the instruction text isn’t directly involved in the reasoning process. Rather, Instructions are how the radiologist would like to see change in the shape and therefore, instructions are the outcome of the reasoning process that we assume would take place in a radiologist’s mind by comparing the desired shape with the current shape. As the correction workflow is user-driven, the stopping criterion is based on the clinician/user judgment that the segmentation quality is satisfactory, which is consistent with the intended human-in-the-loop refinement setting.
Meta-Review
Meta-review #1
- Your recommendation
Provisional Accept
- Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.
Reviewers recommend weak reject, weak accept, and accept. My personal recommendation after skimming: accept (at the very least, for the clever idea: the paper looks at using natural language instructions to fix mistakes in 3D medical segmentation: clearly fills a gap that has been ignored so far. The authors also built a new benchmark, and the work fits well with current trends in VLMs. Despite the provisional accept, the authors are advised to address the concerns of the reviewers (code release, real-world validation, some figure clarity, etc.).
