List of Papers Browse by Subject Areas Author List
Abstract
Temporomandibular Disorders (TMD), affecting the temporomandibular joint and associated musculature, often cause significant functional impairments. Recent advances in deep learning have boosted diagnostic efficiency, yet TMD diagnosis remains challenging due to the requirement of evaluating dual-position (open/closed-mouth) images. Existing models either fail to effectively capture inter-position interactions or rely heavily on expensive, pixel-level annotations as extra guidance. Moreover, they generally overlook the valuable insights offered by raw diagnostic reports, which offer holistic and mixed descriptions of both positions. In this paper, we propose PoCoLT, a novel framework that synergizes Position-aware Cross-modal Learning with LLM-guided hierarchical reports for enhanced TMD diagnosis. Notably, PoCoLT establishes a novel TMD-tailored Position-Modal paradigm that encompasses both representation construction and alignment learning strategies. For Position-aware Intra-modal Construction (PIC), the visual branch employs a cross-position visual fusion (CVF) module to merge position-specific representations into a cross-position one, assigning higher weights to channels with richer global understanding. For the textual branch, holistic raw reports undergo an LLM-guided Textual Decomposition (LTD) process to generate three hierarchical sub-reports, deriving two position-specific and a cross-position textual embedding. Founded on PIC-derived hierarchical representations, PoCoLT devises a Position-aware Cross-modal Alignment (PCA) strategy at corresponding levels. Operated in a supervised manner, PCA progressively shapes a well-clustered textual space and performs precise semantic alignment in the cross-modal joint space. Experiments show that PoCoLT outperforms representative diagnostic and vision-language baselines in practical report-free TMD pre-evaluation.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/3172_paper.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to the Code Repository
N/A
Link to the Dataset(s)
N/A
BibTex
@InProceedings{ZenXin_PoCoLT_MICCAI2026,
author = { Zeng, Xinyi AND Li, Hao AND Zeng, Pinxian AND Luo, Xuwei AND Han, Jize AND Tan, Shuai AND Wang, Yan},
title = { { PoCoLT: Position-Aware Cross-Modal Learning with LLM-Decoupled Hierarchical Reports for TMD Diagnosis } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 16885},
month = {September},
page = {pending}
}
Reviews
Review #1
- Please describe the contribution of the paper
The paper introduces PoCoLT, a framework designed to automate the diagnosis of Temporomandibular Disorders (TMD) by analyzing MRI scans from both open and closed mouth positions. Its primary contribution is the development of a Position-Modal paradigm that organizes diagnostic data into three hierarchical levels: open-mouth, closed-mouth, and cross-position features. This structure is built using two main components: Position-aware Intra-modal Construction (PIC) and Position-aware Cross-modal Alignment (PCA). The PIC component utilizes a large language model to break down general medical reports into three specific sub-reports while using a visual fusion module to integrate features from different mouth positions. The PCA strategy then aligns these visual and textual representations through two training phases—one to refine the text categories and another to match images to those categories using supervised learning. This approach allows the model to use detailed report information during training to improve its diagnostic accuracy even when written reports are unavailable during testing.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
1.It successfully handles the spatial differences between open and closed mouth images by using a specialized visual fusion module that prioritizes the most informative features from both positions. 2.It uses an LLM to extract specific details from general reports, ensuring that the model learns from precise medical descriptions rather than vague summaries. 3.It employs a two-stage supervised learning approach that achieves significantly higher accuracy than existing methods, even when medical reports are unavailable during testing.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.The study relies entirely on a small, private dataset (496 samples) without any external validation. This makes it impossible to know if the model will perform well on diverse patient populations or data from different MRI machines. 2.The framework depends heavily on an LLM (GPT-4o) perfectly parsing text. Because real-world clinical reports are often subjective, messy, or contradictory, any parsing errors will create flawed “ground truth” data that corrupts model training. 3.The model only uses static MRI images and text reports, completely ignoring other vital diagnostic factors that doctors use, such as patient history, pain scales, demographics, and dynamic jaw movement videos.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The paper presents a technically interesting approach to TMD diagnosis with the proposed PoCoLT framework, which effectively utilizes a Position-Modal paradigm to align dual-position MRI scans with LLM-decoupled hierarchical diagnostic reports. This method of using LLMs to extract fine-grained semantic anchors for cross-modal alignment is a meaningful methodological contribution to medical image analysis. However, my concerns are that the model’s clinical feasibility and robustness are constrained by several significant limitations. The empirical validation relies exclusively on a small, private dataset without any external validation, leaving the model’s generalizability completely unproven. Furthermore, the framework’s heavy reliance on the flawless parsing of clinical text by an LLM makes it highly vulnerable to the noisy, subjective, or contradictory nature of real-world medical reports. Lastly, the diagnostic scope is quite narrow, ignoring essential clinical contexts that doctors routinely use, such as patient history, pain scales, and dynamic joint movements. While the architectural design is novel and surpasses the acceptance threshold, the lack of robust external validation and broader clinical integration keeps my score marginal.
- Reviewer confidence
Somewhat confident (2)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #2
- Please describe the contribution of the paper
This paper presents PoCoLT, a vision-language framework for temporomandibular disorder (TMD) diagnosis from dual-position (open/closed-mouth) MRI scans. The key idea is a “Position-Modal paradigm” that constructs hierarchical representations across positions and modalities: on the visual side, a Cross-position Visual Fusion (CVF) module uses bilateral cross-attention and weighted channel attention to merge open- and closed-mouth features; on the textual side, an LLM (GPT-4o) decomposes holistic clinical reports into position-specific and cross-position sub-reports (LTD process). The framework then performs supervised cross-modal alignment (PCA) at each hierarchical level in two phases: first refining the textual space, then aligning visual features to textual anchors. At inference time, only the visual branch is used (report-free), making it practical for real clinical workflows.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
-
The problem formulation is well-motivated. TMD diagnosis genuinely requires analyzing dual-position MRI, and the mismatch between holistic reports and position-specific images is a real challenge. The authors clearly articulate why naive concatenation or standard contrastive learning fall short.
-
The Position-Modal paradigm is a thoughtful design. Decomposing the problem along both the position axis (open, closed, cross-position) and the modality axis (visual, textual) provides a principled structure for the method. The idea of using reports only during training and discarding them at inference is clinically sensible.
-
The CVF module addresses the spatial misalignment issue in a reasonable way. Rather than forcing pixel-level alignment, it operates in feature space with cross-attention and channel-wise reweighting. This is a pragmatic solution.
-
The two-phase training strategy for PCA (first clustering textual embeddings, then aligning visual features to them) is a nice touch. The observation that unsupervised contrastive learning hurts in small-category, small-dataset settings is valid, and the supervised alternative is well-justified.
-
The ablation study is thorough and progressive, clearly showing the contribution of each component (CVF, supervised vs. unsupervised alignment, preliminary phase, LTD vs. regex decomposition).
-
The qualitative results (CAM and t-SNE visualizations) complement the quantitative findings and help build intuition about what the model is learning.
-
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
-
The dataset is private and quite small (248 patients, 496 samples, 3:1 train/test split). With only ~124 test samples and three classes, the reported improvements, while numerically large, are hard to interpret with full confidence. No cross-validation is used, and a single fixed split on such a small dataset raises concerns about generalizability. The authors should consider k-fold cross-validation or at least report confidence intervals beyond the 5-run standard deviation.
-
The reliance on GPT-4o for report decomposition is a practical concern. First, it introduces a dependency on a commercial, non-reproducible API whose version and behavior may change over time. Second, there is no analysis of the quality or failure modes of the LTD decomposition. How often does GPT-4o produce incorrect or hallucinated decompositions? What happens to downstream performance when the decomposition is noisy?
-
The paper does not discuss any failure cases. Where does PoCoLT go wrong? Are there specific TMD subtypes or imaging conditions where the method struggles?
-
The method introduces multiple components, hyperparameters (alpha, training epochs for each phase, embedding dimension), and design choices, but there is limited sensitivity analysis beyond the ablation. For instance, how sensitive is performance to the value of alpha?
-
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
- The paper is generally well-written, though the acronym density is quite high (PoCoLT, PIC, PCA, CVF, LTD, LTD-SR, etc.), which can make it hard to follow on first read.
- It would be helpful to include per-class performance (NDP, DDWR, DDWoR) in the main results, not just aggregate metrics.
- The prompt template used for GPT-4o decomposition should be included (at least in supplementary) for reproducibility.
- How does the method handle cases where the raw report is very short or uninformative?
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The paper addresses a genuine and underexplored clinical problem (dual-position TMD diagnosis) with a well-designed framework. The Position-Modal paradigm is a thoughtful contribution, and the ablation study convincingly demonstrates the value of each component. The improvements over baselines are substantial. However, the evaluation is limited to a small private dataset with a single train/test split, which makes it difficult to confidently assess generalizability. The dependence on GPT-4o for report decomposition raises reproducibility and robustness concerns. The adapted baselines may not fully represent the state of the art. If the authors can address the dataset size concern (e.g., cross-validation results) and provide more analysis on the LTD robustness in a rebuttal, this would strengthen the paper considerably.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Review #3
- Please describe the contribution of the paper
The study proposes a framework that synergizes Position-aware Cross-modal learning with LLM-decoupled hierarchical reports for TMD diagnosis. The framework combines textual inputs derived from LLMs and two visual MRI inputs that are combined using position-aware fusion. This approach is novel in TMD diagnosis.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The paper proposes a novel way to analyse the data already being used in TMD diagnosis in an automated, efficient way. The authors perform strong ablation studies and discuss the results well.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The rationale for the study could be outlined better in the introduction. For readers without any knowledge of TMD it is difficult to understand the problem.
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
Figure 1 contains a lot of information, which becomes not readable without zooming into the document. Maybe the authors could simplify the figure.
“which can be one of the following categories: NDP, DDWR, or DDWoR.” I could not find these categories introduced in the paper.
“For TMD diagnosis, each patient has bilateral temporomandibular joints (TMJs).” This sentence does not make sense; maybe “For TMD diagnosis, the bilateral temporomandibular joints of each patient are analysed/considered”.
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(5) Accept — should be accepted, independent of rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
The paper is well written and organised. The topic is intersting and clinically relevant. A link to the code repository would further strengthen the study.
- Reviewer confidence
Somewhat confident (2)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
N/A
- [Post rebuttal] Please justify your final decision from above.
N/A
Author Feedback
We sincerely thank the Area Chair and all reviewers for their constructive comments and positive evaluation. We are grateful that the reviewers recognized the clinical relevance of the dual-position TMD diagnosis task and our proposed Position-Modal paradigm in the PoCoLT framework. Q1.Report-free inference and offline use of reports/GPT-4o [R1, R2] We clarify that diagnostic reports and GPT-4o are used only as offline preprocessing tools before model training. Raw reports are decomposed into hierarchical sub-reports to provide auxiliary semantic supervision. During inference, neither the LLM nor diagnostic reports are required; PoCoLT relies exclusively on the visual branch to predict TMD categories from dual-position MRIs. This matches clinical pre-evaluation scenarios where reports are unavailable before automated assessment and avoids additional textual input at deployment. Q2.LTD reproducibility and robustness [R1, R2] The LLM does not make clinical decisions or produce diagnostic labels. Its role is limited to offline report structuring. LTD is a constrained extraction task: the prompt specifies the task definition, three sub-report fields, and output schema, so GPT-4o reorganizes existing findings rather than performing open-ended diagnostic reasoning. During dataset curation, cases with missing reports were excluded, and retained samples contain complete paired MRIs and diagnostic reports. The textual branch is used as a training-time semantic anchor: only a lightweight textual projection layer is optimized in the preliminary phase, and the textual branch is frozen during the main phase. To improve reproducibility, the code repository will include the LTD prompt template, parsing schema, and example I/O. The camera-ready will summarize the offline constraints and discuss failure modes for noisy, short, or incomplete reports. Q3.Dataset scale, split, and generalization [R1, R2] We agree that dataset scale and external validation are important. To our knowledge, public TMD datasets with paired open/closed-mouth MRIs, joint-level labels, and matched diagnostic reports remain unavailable. Thus, we constructed a private dataset with 496 TMJ-level samples from 248 patients. To reduce randomness, we conducted five independent runs and reported mean±SD. From the five runs, the run-level 95% CI of ACC is [80.47,81.47] for PoCoLT versus [73.99,76.33] for T3D, supporting robustness under the current split. We acknowledge that five-run SD/CI does not replace split-level or external validation. We will discuss multi-institutional validation across different MRI protocols as an important limitation and future direction. Q4.Clinical scope of PoCoLT and failure patterns [R1, R2] We appreciate the reminder regarding broader clinical contexts. We will revise our claims to position PoCoLT not as a replacement for comprehensive clinical diagnosis, which may involve patient history, pain scales, demographic information, and dynamic jaw movement, but as an automated MRI-based pre-evaluation framework focused on dual-position MRI analysis. According to the confusion matrix, current errors mainly occur between DDWR and DDWoR, where subtle reduction differences require fine-grained dual-position assessment. We will add representative cases in supplementary material. Q5.Clarity, implementation details, and minor revisions [R2, R3] The camera-ready manuscript will define NDP, DDWR, and DDWoR upon first use, revise the bilateral TMJ description, simplify Figure 1, and reduce acronym density. We also clarify that α=0.4 was selected from preliminary validation trials and fixed for all experiments. We thank the reviewers again for their insightful feedback. These revisions will improve clarity, reproducibility, and clinical positioning without changing the method or conclusions.
Meta-Review
Meta-review #1
- Your recommendation
Provisional Accept
- Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.
All reviewers have reached a consensus to accept the paper.
