List of Papers Browse by Subject Areas Author List
Abstract
Multi-modal large language models (MLLMs) have demonstrated increasing promise in medical image understanding, yet their capacity to support clinically aligned diagnostic reasoning for knee joint diseases remains limited because of the scarcity of high-quality datasets that reflect real-world orthopaedic practice. Interpreting knee MRI requires not only recognizing imaging patterns but also structured diagnostic reasoning and severity assessments that reflect orthopedic clinical decision-making. In this work, we introduce \textbf{KJD-25K}, a large-scale, expert-annotated 3D multi-modal dataset specifically tailored for knee joint analysis. It comprises 25K high-quality volume-text pairs with expert-curated diagnostic reasoning and structured severity annotations. %Unlike prior datasets that emphasize anatomical localization, KJD-25K focuses on clinically meaningful disease interpretation and reasoning consistency.
The dataset was developed through over \textbf{4,000} hours of expert clinician effort, ensuring high fidelity to clinical workflows, diagnostic consistency, and semantic richness. Furthermore, we benchmark eight state-of-the-art MLLMs on KJD-25K, establishing comprehensive baselines for clinically grounded knee joint diagnosis. By providing a scalable, clinically validated resource, KJD-25K lays the foundation for developing interpretable, reliable, and workflow-integrated MLLMs for knee joint diagnosis. Data and code will be released upon acceptance.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/0576_paper.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to the Code Repository
N/A
Link to the Dataset(s)
https://huggingface.co/datasets/YiHui0124/KneeCoT
BibTex
@InProceedings{LiYih_KJD25K_MICCAI2026,
author = { Li, Yihui AND He, Along AND Wang, Yuli AND Guo, Kai AND Wu, Yanlin AND Niu, Ben},
title = { { KJD-25K: A 3D Multi-modal Dataset and Benchmark for Knee Joint Diagnosis } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 16895},
month = {September},
page = {pending}
}
Reviews
Review #1
- Please describe the contribution of the paper
First, the authors introduce KJD-25K, a new multimodal dataset comprising knee MRI scans and corresponding radiology reports. The dataset allegedly covers a wide range of pathologies across different knee tissues. Second, the authors benchmark a set of modern open- and closed-source Vision-Language Models on their dataset, and provide baseline performance indicators of the model optimized specifically on KJD-25K.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
1) As excellently summarized in Table 1, public knee MRI datasets are very scarse, and, so far, there has been reported only one dataset specifically designed for VLM-based modeling. Considering that MSK disorders typically involve multiple tissues, a new dataset would be highly beneficial to the community, both for model development and evaluation purposes. 2) The choice of VLMs for evaluation on the proposed dataset is comprehensive and includes multipl state-of-the-art models. Multiple metrics are reported, which provides further insights into potential biases of different models on a standardized benchmark.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
While generally well-written, the paper omits several essential details regarding both the introduced dataset and the evaluation sub-study. This lack of details substantially undermines the potential impact and outreach of the paper. Specifically: 1) Dataset details (language). In Figure 1, the authors mention “Chinese normativity” under “report” block. Are the released radiology reports in Chinese or English? 2) Dataset details (labels). The authors write that the dataset “includes 25’000 . . . volume-text pairs . . . covering a wide spectrum of knee pathologies and severity levels”. While Figure 2 mentions some label categories, a detailed summary of the patologies (along with distributions) would be highly appropriate. Additionally, was severity grading done according to national or international grading schemes? Please, consider elaborating on this aspect in the paper or, at least, on the dataset release webpage. 3) Dataset details (MRI). Next, the authors write “To ensure effective utilization of MRI sequences . . . , we developed an automated procedure to identify and extract the full-coverage scan . . . “. It is unclear exactly which MRI sequences were included in this dataset. Since deep learning models are known to have robustness issues when applied to MRI, providing a description of the included sequences (ideally, also of the scanners and coils) is essential both the dataset presentation and the interpretations of the results in Table 2. 4) Dataset details (inclusion criteria). Next, the authors say “Patient case selection was performed . . . according to predefined inclusion criteria, including image quality, report standardization, and research compatibility. “. What those criteria exactly were? Since demographic and other biases are highly prevalent in medical research, having a detailed description of the criteria is imperative. 5) Benchmark (label derivation). Page 4, paragraph 2: “The label taxonomy was developed . . . and defined through iterative . . . discussions . . . . Using the finalized schema . . . “. What the taxonomy and schema exactly are? They should be explicitly defined. 6) Benchmark (metrics). Page 6, subsection 3.2: “introduce three . . . clinical metrics . . . : lesion detection sensitivity, diagnostic specificity, and structure report compliance . . . “. Neither definition of these metrics nor references to the literature are provided. As these metrics are not “common knowledge”, it is unclear to the reader what is actually being assessed in the model comparison (Table 2) beyond the traditional BLEU and F1 scores. Furthermore, since the dataset labels are also not explicitly presented (see point 2), interpretation of both BLEU and F1 is also questionable.
Lastly, on page 7, the authors state “Fig. 3 qualitatively illustrates the strong generative capability of our model . . . “. According to the taxonomy of scientific evidence (see https://doi. org/10.1097/PRS. 0b013e318219c171), a single case analysis does not constitute high-level evidence. Consequently, describing the generative capability as “strong” based on this qualitative example is unjustified. Please, tone down this conclusion.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
While the authors mentioned that the dataset is “ethically compliant”, it is not clear what the data access permissions for are.
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
-
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(5) Accept — should be accepted, independent of rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
++new dataset +sound design of the benchmark –lack of specific details (hopefully, will be addressed in the released data and source code)
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
Accept
- [Post rebuttal] Please justify your final decision from above.
In their rebuttal, the authors have promised to provide further dataset details in the camera-ready version of the paper and on the dataset release webpage. The lack of such details was one of my primary concerns. Next, I agree with Reviewer 2 on most of their points and concur that the paper might be better-suited for a journal publication. Still, I believe the paper carries sufficient novelty for MICCAI, has the potential to spark constructive discussion at the conference, and can still be extended for a journal publication later on.
Review #2
- Please describe the contribution of the paper
Established KJD-25K, a large-scale 3D multimodal knee joint dataset containing 25000 image and text pairs and structured severity annotations
Built a benchmark by proposed evaluation metrics for clinical perception and linguistics to systematically assess the performance of MLLMs in the diagnosis of knee joint diseases
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
This article constructs the KJD-25K dataset, which is large in scale (with a quantity of over 20000) and focuses on classification, report generation, and Visual Question Answering.
KJD-25K constructs image and text pairs, belonging to the category of Medical Image Computing (MIC), and more precisely, to the current popular multimodal AI.
A benchmark was established based on the KJD-25K dataset to compare the text generation and inference results of different MLLM models on knee joint images. In addition to traditional indicators (BLEU, F1), a comparison of clinical concerns has also been proposed, resulting in the introduction of clinical indicators: LS: Lesion Sensitivity; DS: Diagnostic Specificity; SRC: Structured Report Compliance。
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.The core label construction process of the dataset is opaque. It does not clearly indicate whether these structured labels are re labeled by doctors based on MRI according to a unified protocol, or mainly extracted from reports and then corrected at the case level. The feeling given in the article now is closer to the latter.
2.The literature support for label taxonomy does not hold up. It says’ taxonomy based on priority literature ‘, but the uploaded article [6] is actually a very broad musculoskeletal residue category article in GBD, not the annotation standard for knee MRI, nor is it an imaging grading system for OA/knee joint structural lesions. This reference is not sufficient to support their data label design.
3.The clinical claim is too full. It has always said ‘clinically aligned diagnosis and diagnosis’, but the main focus of the main text is still on report generation/diagnosis, and there is basically no substantial development of diagnosis. This is a bit overrated.
4.Insufficient evidence of clinical effectiveness. The clinical indicators are ultimately evaluated by GPT-4o. However, the reliability, doctor consistency, and labeling protocol of the data labels themselves were not fully explained. This will weaken the persuasiveness of the entire article’s’ clinical reliability ‘.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
This is an eye-catching dataset paper with a timely topic, and I appreciate the effort behind collecting a large-scale knee MRI multimodal dataset and benchmarking multiple MLLMs. I am also interested in seeing this work developed further. However, given the MICCAI page limit (10 pages, with 2 pages already used for references), I feel that many important details are currently missing or under-specified, which makes it difficult to fully assess the validity and reproducibility of the dataset and the benchmark. In particular, the paper would benefit from a much more complete description of the annotation pipeline and label construction process, including whether and how the image findings were assessed in a quantitative or semi-quantitative manner, what exact criteria were used for each structured label, and how expert review was performed in practice. The benchmark section would also be stronger with more ablation studies, especially regarding prompt design and evaluation settings for the MLLMs. Overall, I think this is a promising work with clear potential, but in its current form it feels more suitable for a full-length journal venue where the dataset construction, annotation protocol, and benchmarking details can be presented in sufficient depth. A venue such as npj Digital Medicine or Scientific Data may be a better fit for this type of contribution.
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
It is like an attention grabbing dataset paper, but the current evidence is insufficient to support its core claims about data quality, label validity, and clinical value. That is to say, there is a thematic advantage, but the foundation in methodology and data science is not stable enough. That’s why I lean towards not accepting.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
Reject
- [Post rebuttal] Please justify your final decision from above.
I really admire the meaning of this dataset to the community. However, I cannot accept the details of your dataset only will existing in supplementary and GitHub repository.
Furthermore, the authors not report the population characteristics and statistical result of taxonomy. I also have not idea whether the grading is based on MOAKS or other systems.
Let us omit the R3 who only focuses on MLLM and VQA. The R1 give you accept but raise the most questions, which is also I want to know about the details. But I still think it will be a good paper that submit to a journal so that present more. Actually, I don’t mind whether it eventually accept or not by MICCAI.
Review #3
- Please describe the contribution of the paper
This paper introduces KJD-25K, a large-scale, expert-annotated 3D multi-modal dataset for knee joint diagnosis, comprising 25,000 MRI volume-text pairs with structured severity annotations and expert-curated diagnostic reports. The dataset was constructed through a four-step pipeline involving over 4,000 hours of clinician effort, encompassing expert-guided case selection, multi-modal data structuring (aided by Qwen3-VL for report parsing), AI-assisted label generation (via Qwen3-Max), and rigorous validation. The authors additionally propose a hybrid evaluation protocol combining linguistic metrics (BLEU, F1) with three LLM-based clinical metrics—lesion detection sensitivity (LS), diagnostic specificity (DS), and structured report compliance (SRC)—and benchmark eight state-of-the-art MLLMs (both open-source and closed-source) on KJD-25K. Fine-tuning Qwen3-VL with LoRA on KJD-25K demonstrates consistent improvements across all metrics compared to zero-shot baselines, highlighting the dataset’s utility for clinically grounded knee MRI interpretation.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
KJD-25K is, to the authors’ knowledge, the largest publicly available 3D multi-modal knee MRI dataset to date (25K volume-text pairs). Compared to prior datasets (Table 1), KJD-25K simultaneously provides structured labels, clinical reasoning narratives, and direct VLM trainability—a combination not achieved by predecessors such as OAI, SKM-TEA, or KMAR-50K.
The construction involved five clinicians with 10+ years of experience, seven trained physician assistants, and three clinically trained engineers, totaling approximately 4,000 hours of clinician and GPU effort. This level of expert involvement is commendable and helps ensure clinical fidelity—particularly the four-stage curation pipeline (case selection, data structuring, label generation, validation) that mirrors real-world quality assurance workflows.
The paper provides a thorough benchmark spanning general-purpose open-source MLLMs, closed-source models, and medical-domain models across multiple model sizes. This breadth of comparison offers a useful reference point for the community.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.The data were “collected from a local hospital” (Section 2.1). It would be helpful to briefly discuss generalizability considerations and potential plans for multi-center extension.
2.The paper mentions fine-tuning Qwen3-VL with LoRA but does not specify which variant was used or the training setup. A brief description of the fine-tuning configuration would be appreciated.
3.The paper would benefit from a brief discussion of the benchmark’s current limitations and directions for future improvement, which would help guide community adoption.
4.Limited model diversity in evaluation. The benchmarked models do not sufficiently cover the existing landscape of medical MLLMs. Including sota general medical vision-language models such as GMAI-VL-R1, UniMedVL, and Med-Flamingo would provide a more representative picture of how current approaches perform on this benchmark.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The authors claimed to release the source code and/or dataset upon acceptance of the submission.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
KJD-25K fills a genuine gap as the largest publicly available 3D multi-modal knee MRI dataset with expert-annotated diagnostic reports and structured severity labels.
The comprehensive benchmarking across diverse MLLMs and the clinically motivated evaluation metrics are well-designed. Minor concerns include the single-center data source, missing fine-tuning details, and the lack of a limitations discussion. Overall, the dataset contribution is solid and addresses a real need for the community.
With these additions and clarifications, the paper’s quality should be improved if the concerns are fully addressed in the rebuttal.
- Reviewer confidence
Confident but not absolutely certain (3)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
Accept
- [Post rebuttal] Please justify your final decision from above.
Overall, my concerns has been addressed and I satisfied with the rebuttal. It is a solid work in the medical field for this particular problem and do not forget to revise all the changes in your final revision.
Author Feedback
We sincerely thank the reviewers for their insightful feedback and the value of KJD-25K, the largest publicly available knee MRI report generation dataset and benchmark to date. Due to page limits, additional details and supplementary materials will be hosted in our public repository, and all major concerns will be fully addressed in the camera-ready version. 1.Dataset Transparency and Exclusion Criteria (R #1) All radiology reports are written in Chinese. KJD25K includes four standard clinical knee MRI sequences: T1WI, T2WI, T2WIFS, and PDWIFS, acquired on 3.0T scanners with dedicated 8-channel knee coils. Exclusion criteria included: (1) metallic implants causing severe artifacts; (2) incomplete records or missing reports; (3) acute trauma examinations within 72 hours after injury. 2.Label Taxonomy and Annotation Schema. (R #1, #2) The annotation schema was developed through iterative discussions among five board-certified orthopedic radiologists, following routine knee MRI reporting workflows and has clear literature support. The label type includes: (1) anatomical structures; (2) lesion and pathogenesis categories; (3) structured severity annotations and clinically relevant prediction targets. We agree that [6] alone was insufficient for the taxonomy design; we will add more grading and annotation references from AAOS and COA. 3.Annotation Workflow and Quality Control. (R #2) KJD-25K was built via a four-stage human-AI collaborative pipeline involving a 15-person professional team. AI models were used only to assist preliminary structuring and candidate label generation. All the labels have been retained only after being verified by experts. 4.Evaluation and Metrics. (R #1#2#3) The proposed metrics were adapted from structured criteria routinely used in knee MRI interpretation training,and all these indicators were obtained after discussions with the doctors. LS: proportion of ground-truth lesions correctly identified. DS: proportion of normal structures correctly recognized as normal. SRC: consistency with predefined structured reporting fields. Metrics are scored from 0-100 via structured prompts, with clearer definitions to be added in Section 3.2. We iteratively refined prompts using OpenAI’s GPT-4o on 100 randomly sampled cases, and senior orthopedic radiologists confirmed strong qualitative agreement with expert assessment, supporting its effectiveness as an evaluator. We acknowledge that including more medical MLLMs would further strengthen the benchmark and remains an important future work. 5.Fine-Tuning Details. (R #3) We fine-tuned Qwen3-VL-Instruct-8B using LoRA with rank=8, alpha=16, dropout=0.05, learning rate=2e-4, batch size=16, and 3 training epochs. Both the pre-trained vision encoder and LLM backbone remained frozen during training. 6.Overstated Clinical Claims. (R #1 #2) We agree that several claims were overstated. We replaced “clinically aligned diagnosis” with “clinically aligned radiology report generation” to clarify that the task evaluates coherent report generation rather than autonomous diagnosis. We also removed the unsupported claim of “strong” capability based on a single qualitative case. Furthermore, the study was conducted in close collaboration with clinicians within the hospital, and the experiments were iteratively validated in a clinical setting, highlighting the clinical relevance of our work. 7.Ethics, Reproducibility, and Generalizability (R #1) The IRB-approved, deidentified KJD-25K dataset will be released under CC BY-NC-SA 4.0 via controlled access for approved noncommercial research, consistent with institutional privacy requirements. 8.Multi-center Extension and Benchmark Limitations. (R #3) We acknowledge that the use of a single-center dataset may limit the generalizability of our work. While multi-center falls outside the scope of the current study, it is a highly promising direction for future work. We will discuss limitations and directions for future improvement.
Meta-Review
Meta-review #1
- Your recommendation
Invite for Rebuttal
- Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.
The paper introduces a large multimodal knee MRI-report dataset and benchmarks several vision‑language models
Reviewers agree the dataset is valuable. Public knee MRI datasets for multimodal learning are rare. The large size of the dataset and the wide benchmark are seen as strengths.
However, reviewers raise important concerns. Many dataset details are missing. These include the report language, label taxonomy, severity grading standards, MRI sequences, scanners, and inclusion criteria. The label construction process is unclear, and literature support for the taxonomy is weak. The proposed clinical metrics (LS, DS, SRC) are not clearly defined and rely on LLM based evaluation. Claims about clinical alignment and strong generative ability are seen as overstated.
One reviewer recommends clear acceptance due to the dataset’s potential impact. Other reviewers lean toward weak reject, mainly because of missing methodological and dataset transparency. Reviewers agree the idea is promising, but key details listed above are missing.
Therefore, I recommend rebuttal. In the rebuttal, the authors need to clearly explain the dataset construction, annotation process, label validity, and evaluation methodology.
One reviewer has flagged an ethical concern: “While the authors mentioned that the dataset is “ethically compliant”, it is not clear what the data access permissions for are.”
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Accept
- Please justify your recommendation.
I agree with R2 that this is mainly a dataset paper and therefore needs a more detailed description of the data, annotations, and assessments.
However, I see value in the domain and in the potential use of this dataset for the research community. Therefore, I emphasize that the authors make the dataset, annotations, and all related details fully open to the research community, and carefully revise the camera-ready version according to the reviewers’ comments.
Meta-review #2
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Accept
- Please justify your recommendation.
The authors have provided the requested modifications and improvements. The work and the dataset have important impact in the clinical translation of such technologies.
Meta-review #3
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Accept
- Please justify your recommendation.
I recommend acceptance. The reviewers agree that KJD-25K is a valuable large-scale multimodal knee MRI dataset and benchmark. The main concerns were missing dataset details, unclear label taxonomy and annotation workflow, LLM-based clinical metrics, overstated clinical claims, and ethics/data-access information. The rebuttal addressed several of these issues by clarifying the report language, MRI sequences, exclusion criteria, annotation workflow, metric definitions, fine-tuning setup, moderated claims, and controlled-access IRB-approved data release. Although Reviewer 2 remains concerned that important dataset and taxonomy details are not sufficiently present in the main paper, Reviewers 1 and 4 support acceptance after rebuttal.
