Abstract

Microscopic hyperspectral image (MHSI) provides rich spectral information that enables the detection of subtle biochemical variations in tissue, making it a powerful modality for computational pathology. However, hyperspectral data exhibits substantial spectral redundancy, which may obscure critical discriminative cues. Current approaches rarely focus on identifying informative spectral components, thus failing to leverage the spectral invariance of pathological samples. Moreover, large-scale precise annotations are difficult to obtain, and textual description for MHSI remains largely unavailable, limiting supervised learning. To address these challenges, we propose a novel \textbf{Lo}w-rank \textbf{T}ext-guided \textbf{S}pectral learning \textbf{Net}work for semi-supervised MHSI segmentation, named as \textbf{LoTS-Net}. Specifically, we introduce a text-guided spectral channel selection mechanism to extract refined low-rank spectral representations and incorporate textual semantics to enable interpretable and redundancy-aware channel selection. Then, we leverage intrinsic spectral similarity across samples by constructing a spectral feature container that retrieves correlated prototypes to improve spectral-spatial fusion. This container mitigates noise during the interaction between labeled and unlabeled samples. To support future research, we construct the first microscopic hyperspectral vision-language benchmark and conduct comprehensive evaluations. Experimental results demonstrate that our method significantly outperforms the state-of-the-art methods, particularly under the challenging 5\% labeled data setting. Code is available at \url{https://github.com/ECNU-MultiDimLab/LoTS-Net}.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/4532_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

https://github.com/ECNU-MultiDimLab/LoTS-Net

Link to the Dataset(s)

N/A

BibTex

@InProceedings{ZhaSiq_LowRank_MICCAI2026,
        author = { Zhang, Siqi AND Zhang, Qing AND Wang, Yan AND Li, Qingli},
        title = { { Low-Rank Text-Guided Spectral Learning for Semi-supervised MHSI Segmentation } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16880},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    They constructed a new semi-supervised MHSI segmentation method with spatial and spectral text guides.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    The method selects useful spectral channels and enhances spatial features using features extracted from spatial and spectral texts. In addition, I think the spectral feature container is also a strength of the paper. Since it plays the role of a spectral dictionary, the model can do segmentation using spectral characteristics.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    In the experiment, spatial and spectral texts are generated from MHSI, so I wonder where the generated text can truly reflect the spatial and spectral characteristics of cancers and abnormal regions. When submitting a journal, it would be even better if there were expert reviews of the texts.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (6) Strong Accept — must be accepted due to excellence

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    This paper is well-written, and the approach and the results have an impact in this field.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    N/A

  • [Post rebuttal] Please justify your final decision from above.

    N/A



Review #2

  • Please describe the contribution of the paper

    The paper introduces a low-rank text-guided spectral learning network (LoTS-Net) for the semi-supervised segmentation of microscopic hyperspectral images (MHSI). The authors extend two existing image datasets with text prompts to produce a novel benchmark and, using those, show superior performance of their approach over the state-of-the-art.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    • The work is pioneering in formulating the task of text-guided semi-supervised segmentation for microscopic hyperspectral images. A similar topic has been explored before for medical image segmentation, but not for MHSI.
    • The proposed approach is extensively evaluated to illustrate its superiority for the task.
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
    • While the contribution itself is very interesting and seems valuable, the text is unclear and hard to follow at places. For example, Sections 2.2 and 2.3 are unclear and fail to introduce some of the notation.
    • The authors promise to release the code. For reproducibility purposes, it would also be helpful if the datasets (or, rather, the corresponding text prompts) were released too.
  • Please rate the clarity and organization of this paper

    Poor

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
    • The paper needs careful proofreading to fix misspelling, semantic and styling errors, etc.
    • Fig. 1 is practically illegible on A4/letter paper; it needs to be bigger, especially the textual parts.
  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper is ‘borderline’: overall, the work is very promising, but its presentation seems raw.

  • Reviewer confidence

    Somewhat confident (2)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The contribution is not very strong, but I think it is acceptable in the revised form, since it explores a novel conceptual idea.



Review #3

  • Please describe the contribution of the paper

    The paper proposes LoTS-Net, a semi-supervised framework for microscopic hyperspectral image segmentation that combines text-guided spectral channel selection, a spectral feature container for retrieval-based spatial-spectral fusion, and consistency-based training on unlabeled data. Its main contribution is to incorporate language guidance into spectral feature selection and semi-supervised MHSI segmentation, while also extending two public datasets with text prompts to support this multimodal setting.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1.The paper studies an interesting and relatively underexplored problem, namely semi-supervised microscopic hyperspectral image segmentation with language guidance, and the combination of spectral modeling and text guidance is well motivated for this setting. 2.The empirical evaluation is fairly solid. The paper compares against multiple relevant fully supervised and semi-supervised baselines, including close related methods such as ACCL-CINet, Spec-Tr, Omni-Fuse, and several text-guided segmentation models, and it also includes component ablations.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1.The paper combines several recognizable components, including text-guided feature enhancement, spectral channel selection, retrieval from a prototype/container memory, pseudo-labeling, and consistency regularization. This is a reasonable system design, but it is less clear which element constitutes the main conceptual advance. As a result, the contribution feels more like a careful combination of existing techniques tailored to MHSI segmentation than a clearly new methodological direction. 2.Low-rank structure is emphasized repeatedly in the title, abstract, and motivation, suggesting that explicit low-rank modeling is a central technical contribution. However, from the method description, the implementation appears to rely primarily on learned spectral channel selection and text-guided filtering rather than on a clearly defined low-rank decomposition, low-rank constraint, or explicit rank-regularized objective. This makes the low-rank framing feel somewhat stronger than what is directly supported by the actual formulation. 3.While the paper includes an ablation study, it remains relatively coarse and does not analyze key hyperparameters or submodule choices. In particular, there is no sensitivity study for the number of selected spectral channels, prompt design, text encoder choice, or container design, which makes it harder to understand which design decisions are essential to the reported gains.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper addresses an interesting problem and includes a better-than-average empirical comparison against several relevant baselines, which is a real strength. However, I am not yet fully convinced that the methodological contribution is sufficiently clear and strong for acceptance. In particular, the paper’s framing around “low-rank” spectral learning is not fully matched by the actual formulation, and the ablation analysis is not deep enough to isolate the importance of several key design choices. Overall, the work is promising and technically competent, but in its current form it feels somewhat incremental and under-justified relative to the paper’s claims.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The previous review addressed my concerns.



Author Feedback

We thank the reviewers and AC. We pioneer a semi-supervised MHSI framework exploiting dataset-wise spectral consistency, resolving spatiospectral redundancy to extract low-rank features.

@R3: conceptual advance Our method is a paradigm shift specific to MHSI, not a simple combination of existing modules. Our core advance lies in leveraging the physical prior that the same pathological tissues share invariant intrinsic spectral signatures across different samples. By using a text-guided spectral selector and a global queue, unlabeled samples match their spatial features with a pure, historical spectral dictionary, resolving the error-accumulation bottleneck in MHSI SSL.

@R3: low-rank formulation Methods like low-rank decomposition are designed for key spectral feature extraction. Aligning with the principle, we first employ DwConv and Non-negative Matrix Factorization in spectral encoders to extract low-rank spectral features from individual samples. Subsequently, guided by texts with key absorption indices and lesion descriptions, we further eliminate residual redundant spectral and spatial information. Most importantly, to exploit the cross-sample key spectral consistency inherent in MHSI datasets, we store these purified features in a global queue and fuse them with relevant spatial features, elevating traditional intra-sample low-rank extraction to dataset-level.

@R3: ablation studies on 20% labeled MDC and GPCC 3.1) Number of selected channels K: Testing K∈{5,10,15,20} yields DSC scores of {78.69,80.18,81.37,80.90}% on MDC and {88.91,89.85,90.59,89.29}% on GPCC. K=15 optimally balances information retention and noise reduction. 3.2) Prompt design: We kept the architecture fixed and altered the text inputs. a. Spatial-only: Feeding only spatial texts to both streams leads to performance decrease, with DSC of 80.46% on MDC and 89.66% on GPCC. b. Merged: Feeding the concatenated spatial-spectral texts to both streams yields 79.88% and 88.60%. Our original design achieves the highest DSC 81.37% and 90.59%, proving that specific modalities strictly require corresponding semantic guidance. 3.3) Text encoder: We replace the biobert-large with a general-domain bert-large-cased and a smaller biobert-base. These variants yield DSCs of 79.84% and 80.22% on MDC, and 89.98% and 90.25% on GPCC, indicating strong robustness of our model to varying text inputs. 3.4) Length of container: Varying the queue length N_Q∈{450,600,750,900} results in DSCs of {79.78,81.00,81.37,80.84}% on MDC and {88.89,89.84,90.59,90.29}% on GPCC. N_Q=750 provides sufficient historical anchors, avoiding most stale noise.

@R2: details 4.1) Notations: a. In 2.2, we revise the text into: “For Text-Guided Spectral Channel Selection, spectral text prompt T_{spec} is encoded by a frozen text encoder to obtain the semantic feature F_{spec,t}, which is injected into F_{spec,s}^\prime through a cross-attention mechanism, producing an enhanced spectral feature F_{spec}^\pprime. F_{spec}^\pprime is then projected into importance scores for each channel. Then the top-K channels are selected.” b. To clarify Eq.(1) In 2.3, we insert: “where d represents the feature dimension of {\hat{F}}_{spa}, {\mathrm{argmax}}^{(-1)} denotes the argmax operation along the last dimension” right after the equation. 4.2) We correct “spatiaospectral” to “spatiospectral” in 2.4 and “amoun” to “amount” in 3.2.We revise instances including “a enhanced” to “an enhanced”, “refined feature… serve” to “serves” in 2.2 and “strategy… improve” to “improves” in 3.3.4.3) In the camera-ready version, we will enlarge Fig. 1, particularly the textual parts and tensor shapes. 4.4) For complete reproducibility, we commit to open-sourcing the dataset, the code and pre-trained weights.

@R1: text authenticity All generated text prompts were already verified by professional pathologists, ensuring they truly reflect pathological characteristics of cancers.




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    This submission proposes hyperspectral image (MHSI) segmentation using text guidance. The reviews are found constructive and converges on the high novelty of the proposed application on MHSI, text-prompt extensions to two public datasets, and solid evaluation. Possibly fixable issues were noted on a missing support for the motivation and contribution of using low-rank modeling, as well as clarity of the methodology and figures. Addressing the low-rank contribution would improve the appreciation of the submission. For all these reasons, and with respect to the other submissions, the recommendation is requesting a rebuttal.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The rebuttal was comprehnensive and has clarified the low-rank formulation with detailed ablation studies. The reviewers reached a consensus to accept this framework for text-guided semi-supervised microscopic hyperspectral image segmentation.



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    After the rebuttal, all reviewers agreed that the paper should be accepted. The Area Chair supports this consensus and recommends acceptance.



Meta-review #3

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The proposed low-rank text-guided spectral learning network (LoTS-Net) is recognized as a pioneering and well-motivated framework for semi-supervised microscopic hyperspectral image segmentation. While initial concerns were raised regarding the clarity of the low-rank formulation, the clinical validity of the generated text, and presentation issues, consensus among reviewers post-rebuttal highlights that the authors have successfully resolved these queries through detailed clarifications and committed revisions.



back to top