List of Papers Browse by Subject Areas Author List
Abstract
Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains challenging due to low contrast, frequent discontinuities, and severe class imbalance. Although recent convolutional and Transformer-based models have improved performance, they often yield fragmented predictions and fail to recover fine branches. We propose CSWinUNETR, a task-oriented 2D/3D backbone for thin-structure segmentation. It employs cross-shaped stripe self-attention to model long-range principal-axis context and incorporates cyclic shifts to enhance information exchange across stripes. To better preserve fine-grained details, we further introduce a detail-enhanced multi-scale self-attention module that aggregates contextual features from multi-resolution representations. In addition, we propose sparse-control dynamic snake convolution, which reconstructs reliable dense curvilinear kernels from sparsely predicted control points to better follow tortuous geometry. Extensive experiments on four benchmarks across ophthalmology, neurovascular imaging, and dermatology demonstrate that CSWinUNETR consistently outperforms state-of-the-art methods without task-specific post-processing or topology-aware losses. Code is available at https://github.com/labhai/CSWinUNETR.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/4831_paper.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to the Code Repository
https://github.com/labhai/CSWinUNETR
Link to the Dataset(s)
N/A
BibTex
@InProceedings{MooJun_CSWinUNETR_MICCAI2026,
author = { Moon, Junho AND Chung, Haejun AND Jang, Ikbeom},
title = { { CSWinUNETR: Segmentation of Thin Anatomical Structures in Medical Images } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 16883},
month = {September},
page = {pending}
}
Reviews
Review #1
- Please describe the contribution of the paper
The paper’s main contribution is the introduction of CSWinUNETR, a unified backbone for segmenting thin and tortuous anatomical structures in both 2D and 3D medical images. The method combines three key design elements: shifted cross-shaped window self-attention for orientation-aware long-range dependency modeling, a detail-enhanced multi-scale self-attention module for preserving fine structural cues, and sparse-control dynamic snake convolution for more stable geometry-aligned local aggregation along curvilinear structures. Overall, the paper’s contribution lies less in proposing an entirely new segmentation paradigm and more in presenting a targeted architectural integration for thin-structure segmentation, together with empirical validation across multiple heterogeneous benchmarks.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The paper has several notable strengths. First, it addresses an important and challenging problem, namely the segmentation of thin and tortuous anatomical structures, where preserving fine branches and structural continuity is particularly difficult. Second, the method is technically well motivated: while the individual ingredients are not entirely new on their own, the paper combines cross-shaped attention, multi-scale detail enhancement, and a sparse-control snake-style convolution in a coherent way that is well aligned with the target problem. In particular, the sparse-control formulation is an interesting refinement aimed at improving stability when modeling curvilinear trajectories. Third, the evaluation is relatively comprehensive, covering four heterogeneous datasets across both 2D and 3D settings, with quantitative, qualitative, and ablation results that make the empirical case reasonably convincing.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The methodological novelty appears moderate rather than substantial. 1.The proposed method combines several design elements that are already established in prior work at the component level. In particular, the use of cross-shaped stripe attention is closely related to CSWin Transformer: A General Vision Transformer Backbone with Cross-Shaped Windows (2022), while the cyclic shift mechanism follows the design idea popularized by Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows (2021). In addition, the use of snake-style dynamic convolution for curvilinear structure modeling is strongly related to Dynamic Snake Convolution Based on Topological Geometric Constraints for Tubular Structure Segmentation (2023). Moreover, CSWin-UNet: Transformer UNet with Cross-Shaped Windows for Medical Image Segmentation (2025) has already adapted CSWin-style attention to medical image segmentation. Therefore, the main contribution of the present paper seems to lie more in the integration and task-specific refinement of existing ideas than in a fundamentally new segmentation formulation.
2.The efficiency and deployment trade-offs are not characterized in sufficient detail. Although the paper reports parameter counts, it does not provide a more complete efficiency analysis, such as FLOPs, inference speed, or memory usage. This omission is particularly relevant because the method targets both 2D and 3D medical image segmentation and combines attention-based modeling with dynamic curvilinear aggregation, both of which may introduce nontrivial computational overhead. A more thorough efficiency analysis would make the practical utility of the method easier to assess.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
I didn’t give this paper a higher rating because its overall methodological foundation is weak.
It is a careful integration and improvement of existing ideas, refined for a specific task.
Furthermore, while the number of parameters reported is useful, there is limited discussion on trade-offs regarding runtime, memory, or deployment, which are crucial for a comprehensive assessment of practical impact, especially in a 3D environment.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #2
- Please describe the contribution of the paper
The paper proposes CSWinUNETR, a general-purpose 2D/3D segmentation backbone tailored for thin and tortuous anatomical structures. The method integrates three components: (i) cross-shaped stripe self-attention with a cyclic shift to enable efficient long-range, orientation-aware context aggregation; (ii) a detail-enhanced multi-scale multi-head self-attention (MS-MHSA) that fuses multi-resolution features while preserving high-frequency details; and (iii) a sparse-control dynamic snake convolution (SDSConv) that constructs dense, curvilinear sampling kernels from a small set of control points to follow tortuous geometry without cumulative offset drift. Experiments on four benchmarks (FIVES, FFHQ-Wrinkle, TopCoW MRA/CTA) show consistent improvements over strong CNN/Transformer backbones.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
1.The SDSConv module is an interesting refinement of prior dynamic snake/deformable operators. Integrating CSWin attention with a Swin-style cyclic shift is a simple but effective architecture. 2.Evaluation spans four benchmarks across different modalities and dimensions, improving generality. 3.Comparisons with many baselines, including strong backbones (nnUNetv2, UNETR, SwinUNETR/v2), thin-structure-oriented methods (DSCNet, CS²-Net, ER-Net), and nnUNet variants (clDice, cbDice, skeleton-recall).
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.The method section needs more detail. Specifically, the 3D version of SDSConv is not clear. The text mentions two axes (x and y), but 3D images have three (x, y, and z). Please explain how the model handles the third axis in practice.
2.The introduction does not clearly explain the difference or the motivation for your method compared to baselines like CSWin-UNet. Even though you used them as baselines, it is not clear what makes your specific design unique.
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
Overall, I find the work technically sound, relevant, and practically valuable for the MICCAI community. The main areas to strengthen are clarity and completeness in motivation compared with related work and details in SDSConv.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
Accept
- [Post rebuttal] Please justify your final decision from above.
The authors have addressed all of my concerns. Therefore, I recommend accepting the manuscript.
Review #3
- Please describe the contribution of the paper
This paper addresses the challenge of accurately segmenting thin, tortuous anatomical structures in medical images and proposes CSWinUNETR, a general-purpose backbone for 2D and 3D thin-structure segmentation. It employs cross-shaped stripe self-attention to model long-range principal-axis context and incorporates cyclic shifts to enhance information exchange across stripes.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The experiments in this paper are relatively thorough, but the novelty is insufficient.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The novelty of this paper is limited, as the proposed method largely assembles existing research, such as references [5], [7], [19], and [22], without introducing substantial new ideas. Moreover, the experimental results show that, despite a significant increase in the number of model parameters compared to other methods, the proposed approach does not achieve a particularly notable improvement in performance metrics.
- Please rate the clarity and organization of this paper
Poor
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The paper lacks sufficient novelty and the marginal performance gains do not justify its significant increase in model parameters.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
Reject
- [Post rebuttal] Please justify your final decision from above.
As pointed out by both myself and other reviewers, the proposed method appears to be mainly a combination of several existing components derived from prior works, such as references [5], [7], [19], and [22]. However, in the rebuttal, the authors still failed to clearly explain the fundamental differences between their method and these existing approaches. As a result, the actual methodological novelty of the paper remains unclear. In addition, the rebuttal did not provide sufficient evidence to demonstrate that the reported performance improvements genuinely arise from the proposed methodology itself, rather than from the increased model capacity and additional parameters introduced by integrating multiple modules. Given these unresolved concerns and the current limitations of the manuscript, I do not believe the paper is yet strong enough for acceptance at MICCAI.
Author Feedback
We thank the reviewers and AC. We appreciate the recognition that CSWinUNETR is well-motivated for a challenging problem (R1), technically sound and practically valuable (R2), with thorough experiments (R1-3). [R1-3] Novelty Thin-structure segmentation fails in two specific ways: (i) fragmented continuity along curvilinear paths and (ii) loss of fine-scale detail under downsampling. Our contributions are uniquely designed to address these: 1) SDSConv targets (i) via novel sparse-control, stable curvilinear aggregation; 2) detail-enhanced MS-MHSA targets (ii) via spatially adaptive high-frequency and multi-scale feature fusion; and 3) CSWinUNETR provides a 2D/3D backbone that lets them compose, making it a task-specific advancement over CSWin-UNet. 1) & 2) are newly introduced to directly address these failure modes, while 3) couples them with CSWin-attention and cyclic shifts to model orientation-aware context. 1) As R1/R2 acknowledged, SDSConv differs from DSConv by reformulating curvilinear kernel construction from stepwise offset tracking to control-point-based trajectory modeling. It predicts sparse control points and builds kernels via non-cumulative trajectory parameterization. DSConv sequentially accumulates sampling-path offsets, making trajectories sensitive to local prediction errors. This reformulation is particularly beneficial for thin structures, where local evidence is often discontinuous/corrupted by artifacts. 2) Detail-enhanced MS-MHSA forms attention K/V from a spatial fusion of an explicit high-frequency detail branch and multi-scale branches, enhancing boundary/fine-grained structural cues before stripe attention. CSWin-UNet projects K/V from the same feature map without any detail/context conditioning. The spatially adaptive, detail-aware fusion also differs from the fixed-branch concatenation in IncepFormer. 3) CSWin-attention with cyclic shifts bridges weak/interrupted thin-structure evidence via axis-wise propagation and cross-stripe exchange. As R2 noted, this is simple yet effective and mitigates boundary fragmentation without additional parameters. Unlike CSWin-UNet, primarily designed for 2D slice-wise segmentation, CSWinUNETR provides a 2D/3D backbone, enhancing volumetric curvilinear continuity. [R1/R3] Efficiency To clarify deployment cost, we note that parameter count alone does not fully capture inference cost, particularly in the 3D setting. We compare CSWinUNETR (81M) with DSCNet (8M), a parameter-light yet strong-performing model, and SwinUNETRv2 (73M), a parameter-comparable backbone. Under identical inference settings on TopCoW-MRA (input size: 96^3), DSCNet/SwinUNETRv2/CSWinUNETR require 4.23/1.51/1.35 GB memory and 425/346/346 GFLOPs, with per-patch latencies of 321/250/292 ms. CSWinUNETR thus uses low memory with comparable FLOPs/latency while improving performance. [R3] Gains beyond parameter count The reported gains are statistically significant (p<0.05) across most metrics compared with SOTA baselines and stem from module design rather than parameter scale. On TopCoW-MRA, CSWinUNETR outperforms nnUNetv2+clDice by 28.9/14.7/2.5/2.1% on HD95/β/Dice/clDice. Crucially, the ablation in Tab.3 shows that our three contributions add only 0.9M parameters yet yield consistent stepwise improvements across all metrics (30.2/24.7/2.3/1.8% on TopCoW-MRA), directly demonstrating that gains arise from architectural design rather than capacity. [R2] 3D formulation For 3D inputs, CSWin-attention uses x/y/z-stripe orientations with cyclic shifts over all spatial axes. SDSConv consists of three axial branches; each branch traverses one principal axis and predicts deformations along the two orthogonal axes. The predicted control offsets are expanded into dense trajectories, used to sample the volumetric features, and the sampled features are aggregated via 3D convolutions. We will revise §2.1–2.2 for clarity, including the 3D formulation, and incorporate the above clarifications in the final version.
Meta-Review
Meta-review #1
- Your recommendation
Invite for Rebuttal
- Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.
The reviewers appreciated the challenge of segmenting thin and tortuous anatomical structures, and found your proposed CSWinUNETR architecture to be well-motivated with a comprehensive empirical evaluation across multiple 2D and 3D datasets.
Please ensure your rebuttal addresses the following critical issues: 1) All three reviewers expressed concerns regarding the methodological novelty, noting that the architecture relies heavily on assembling existing components . Please clearly articulate what fundamentally distinguishes CSWinUNETR from existing works, specifically baselines like CSWin-UNet. 2) Emphasize the unique, task-specific refinements, such as the sparse-control formulation of the snake convolution, that make this integration more than just a simple assembly of parts. Explain why this specific combination is uniquely suited for thin-structure segmentation. 3) Please clarify how the model handles the third spatial axis ($z$-axis) in practice, given that the text currently only describes $x$ and $y$ axes. Provide clear mathematical or architectural explanations for the 3D extension. 4) Reviewers 1 and 3 raised concerns about the computational overhead and efficiency of the model. Reviewer 3 explicitly stated that the marginal performance gains do not seem to justify the increase in model parameters.
Please respond to the other points raised if space allows in the rebuttal.
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Accept
- Please justify your recommendation.
After considering the reviews and the authors’ rebuttal, I recommend acceptance. The paper addresses a relevant and technically challenging problem: reliable segmentation of thin, tortuous anatomical structures across heterogeneous 2D and 3D medical imaging settings. While the individual components are related to existing CSWin, Swin/UNETR, and dynamic snake convolution ideas, the paper presents a coherent task-specific architecture with clear motivation for thin-structure continuity and fine-branch preservation.
The main strength is the combination of orientation-aware CSWin attention, detail-enhanced multi-scale attention, and sparse-control dynamic snake convolution within a unified 2D/3D backbone. The empirical validation is also relatively strong, covering multiple datasets, modalities, dimensions, topology-aware metrics, qualitative results, and ablations. The rebuttal further clarifies the 3D formulation, the difference from CSWin-UNet and DSConv, and the efficiency trade-off, which addresses several important concerns raised during review.
For the camera-ready version, the authors should clearly clarify the relationship and technical distinctions between CSWinUNETR and closely related components such as CSWin-UNet and DSConv, incorporate the 3D SDSConv formulation and efficiency analysis already clarified in the rebuttal, and moderate overly broad claims about general-purpose superiority where the performance gains are relatively small.
Meta-review #2
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Reject
- Please justify your recommendation.
The main concern across the reviews is that the methodological novelty appears moderate. Reviewers 1 and 3 both viewed the approach as primarily an integration and refinement of existing ideas, including CSWin/Swin-style attention mechanisms and snake/deformable convolutional modeling for curvilinear structures. Overall, the submission is well motivated and has a meaningful empirical component, but the evidence is mixed.
Meta-review #3
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Accept
- Please justify your recommendation.
The paper addresses a challenging task (thin/tortuous structure segmentation) with a well-integrated architecture, and the rebuttal convincingly clarifies the methodological differences from prior work and provides missing efficiency analysis (FLOPs, memory, latency). The reviewers’ concerns about novelty and computational cost were adequately resolved, and the ablation studies confirm that performance gains come from architectural design rather than parameter scaling.
