Abstract

While Large Vision-Language Models (VLMs) demonstrate remarkable generic capabilities, their clinical reasoning in specialized domains like ocular surface diseases (OSDs) is severely hindered by a paucity of high-fidelity, multimodal instruction-tuning data. To dismantle this data bottleneck, we introduce IRIS, an Intelligent Recognition and Interaction System tailored for fine-grained OSD understanding via external eye photography. First, we curate IRIS-120K, the largest and most comprehensive OSD visual question-answering (VQA) dataset to date. Crucially, to overcome the semantic shallowness of conventional image-caption pairs, we propose a synergistic data generation paradigm to explicitly inject clinical priors. Our data engine operates via a dual-branch framework: 1) a Topic Finding Tree (TFT) that hierarchically anchors visual features to precise anatomical and pathological concepts, enforcing rigorous medical deduction logic; and 2) a Scene-driven strategy that synthesizes role-adaptive clinical dialogues to ensure pragmatic generalization. By explicitly aligning a compact 4B-parameter VLM on this structurally enriched corpus, IRIS achieves highly competitive performance and outperforms both generalist and specialized medical VLMs with up to 34B parameters within this dataset. Our findings underscore that structured knowledge injection effectively bridges the domain gap, unlocking the potential for resource-efficient, expert-level AI deployment on mobile edge devices for scalable OSD screening. Code, datasets, and model weights will be publicly released by this \href{https://github.com/hwei-hw/IRIS}{repo}.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/5330_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

https://github.com/hwei-hw/IRIS

Link to the Dataset(s)

IRIS-120K: https://huggingface.co/datasets/hw-hwei/IRIS-120K

BibTex

@InProceedings{WeiHao_IRIS_MICCAI2026,
        author = { Wei, Hao AND Qi, Wenjin AND Dai, Dasen AND Zhang, Minqing AND Yuan, Wu},
        title = { { IRIS: An Intelligent Vision-Language System for Ocular Surface Diseases via Topic Tree and Scene-Driven VQA Generation } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16878},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    This paper presents IRIS, a vision-language system for ocular surface disease understanding from external eye photographs. The work proposed a data-centric framework that constructs IRIS-120K, which the authors position as the largest external-eye VQA dataset to date, using a dual-branch data engine: a Topic Finding Tree (TFT) that organizes supervision around ocular anatomy and pathological findings, and a Scene-driven generation strategy that simulates role-specific clinical interactions. A compact Qwen3-VL model fine-tuned with LoRA on this corpus is reported to outperform substantially larger general and medical VLM baselines across several VQA settings.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    The paper addresses an under-served but clinically relevant modality. Most ophthalmic foundation-model work is concentrated on fundus/OCT, whereas this paper focuses on external eye images for ocular surface disease understanding, which is a meaningful and relatively neglected direction.

    The empirical gains are strong, especially relative to model size.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    The clinical validity of the dataset and test set is not yet sufficiently convincing. A central concern is that much of the pipeline relies on VLM/LLM-based generation and filtering, including the final quality control of the test pool with GPT-5-mini. The paper does not clearly report how much of the final dataset, especially the so-called “gold-standard” test set, was verified by ophthalmologists, nor does it provide inter-rater agreement or systematic error analysis. This weakens confidence that the reported gains reflect clinically reliable improvement rather than better alignment to generated supervision.

    The paper’s novelty is stronger on the data/supervision side than on the model/method side. The backbone appears to remain largely standard Qwen3-VL with LoRA fine-tuning, so the main innovation is the dual-branch data engine rather than a new learning method. That is acceptable for a data-centric paper, but the manuscript currently makes fairly strong claims about methodological advance.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The current version does not yet provide enough evidence for the clinical reliability, reproducibility, and deployment-level claims it makes.

    Another issue is the limited expert-backed validation of the dataset/test set and the lack of clinically aligned evaluation for open-ended responses.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The author response can address my concerns.



Review #2

  • Please describe the contribution of the paper

    This paper introduces IRIS, a vision-language system for ocular surface disease (OSD) diagnosis from external eye photography. The core contribution is IRIS-120K, a large-scale VQA dataset generated through a dual-branch data engine combining a Topic Finding Tree (TFT) — which hierarchically anchors visual features to anatomical regions and pathological findings — and a Scene-driven strategy that simulates role-adaptive clinical dialogues across three user types. A compact 4B-parameter VLM is fine-tuned on this dataset and reported to outperform general and medical baselines up to 34B parameters across closed-ended and open-ended question types.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    S1.Addresses a real clinical gap. External eye photography for OSD screening is a genuinely underserved modality in medical AI, and the motivation for a dedicated dataset and model is well-supported by the documented bias toward posterior segment imaging in existing ophthalmic foundations models. S2.Clinically motivated data design. The TFT’s hierarchical anatomical decomposition and the four-stage reasoning chain (visual observation → clinical correlation → logical deduction → conclusion) are meaningfully aligned with ophthalmological diagnostic practice, going beyond flat caption-based supervision. S3.Consistent cross-scale performance gains. The reported improvements over substantially larger baselines across multiple question types provide a clear empirical signal that domain-specific structured data curation yields gains beyond parameter scaling in this narrow domain.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    W1.Data quality relies heavily on unvalidated LLM-generated content. The entire IRIS-120K dataset — including the clinical descriptions, TFT-derived question-answer pairs, reasoning chains, and scene dialogues — is generated by LLMs (primarily Qwen3-VL-32B) from image-caption pairs of varying quality. The quality control pipeline uses GPT-5-mini to filter hallucinations, but LLM-based filtering cannot reliably detect clinically incorrect but linguistically plausible statements.

    W2.IRIS-120K’s test set is generated by the same data engine and using the same LLMs (Qwen3-VL-32B for descriptions, GPT-5-mini for quality control) that inform the training data. IRIS-4B is then fine-tuned on the training split and evaluated on the test split of this same dataset. This creates a fundamental circularity: the model is being evaluated on data whose distribution it was explicitly optimized to match, using question formats, reasoning structures, and vocabulary that were baked into training. Near-perfect scores on closed-ended tasks (97.25 on Judge, 98.52 on Single-choice) are more plausibly explained by distributional overlap between train and test than by genuine clinical reasoning capability.

    W3.Use of BLEU / ROUGE for clinical VQA and reasoning quality is limited. These metrics often correlate poorly with factual correctness or diagnostic usefulness. Stronger medical evaluation (expert grading, semantic correctness, safety scoring) is needed.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    This paper addresses an important medical AI application and presents a potentially useful specialized dataset plus lightweight model for ocular surface disease understanding. However, the current submission relies heavily on synthetic generated supervision, uses an evaluation setup that may favor the proposed pipeline, and lacks strong external clinical validation. Several claims around interpretability and superiority over much larger models appear stronger than the current evidence supports.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    I believe overall, the work is a valuable contribution and should be considered for acceptance.



Review #3

  • Please describe the contribution of the paper

    This paper introduces IRIS, a vision-language system designed for fine-grained understanding of ocular surface diseases (OSDs) from external eye photography. The authors argue that existing VLMs, both general-purpose and medical, perform poorly on OSD diagnosis due to a severe data gap: current ophthalmic AI is overwhelmingly biased toward posterior segment (retinal) imaging, leaving external eye pathology underserved.

    To address this, the authors curate IRIS-120K, claimed to be the largest OSD visual question-answering dataset (117.7K training pairs, 8.2K test). Data is aggregated from three heterogeneous sources: medical textbooks (3.1K cases), web/network data (93.5K from WeChat articles, PubMed, Google Lens), and public datasets (35.4K). A multi-stage preprocessing pipeline uses existing VLMs (Qwen2.5-VL-7B/72B) and YOLOv7 for filtering, sub-figure segmentation, and caption alignment.

    The core methodological contribution is a “Clinically-driven Data Engine” with two VQA generation branches: (1) a Topic Finding Tree (TFT) that hierarchically decomposes images into 10 predefined ocular anatomical regions and extracts region-specific clinical findings, generating structured 4-stage reasoning chains (Visual Observation → Clinical Correlation → Logical Deduction → Conclusion); and (2) a Scene-Driven strategy that simulates role-adaptive clinical dialogues for three user types (Patient, Doctor, Student) across 12 clinical scenarios. A quality-aware dynamic sampling protocol prioritizes high-fidelity sources during generation, yielding ~130K initial pairs that are deduplicated and filtered (using GPT-5-mini) to produce the final dataset.

    The authors fine-tune Qwen3-VL models (2B, 4B, 8B) via LoRA on 4× A6000 GPUs. Evaluation is performed on the IRIS-120K test set against 16 baselines spanning general VLMs (GPT-5-mini, InternVL-3.5, Qwen3-VL series) and medical VLMs (HuatuoGPT-Vision, HuluMed, Lingshu, Med-Gemma). Metrics include accuracy for closed-ended tasks and BLEU-1/ROUGE for open-ended tasks. IRIS-4B achieves the highest overall average score (74.26), outperforming Lingshu-32B (55.00) by 19.26 points and Qwen3-VL-32B (50.61) by 23.65 points. Ablation studies show TFT and Scene-driven data are complementary, and that performance peaks at 4B parameters.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1) addressing a genuine clinical gap: The paper arguest hat existing ophthalmic AI is heavily biased toward posterior segment imaging, citing that external eye photographs constitute only 0.5% of the EyeCLIP dataset. OSD is a major global health burden (5th leading cause of blindness, affecting up to 34% of adults), and the focus on smartphone-based screening for resource-limited settings is clinically meaningful and timely.

    2) well-designed data curation pipeline: The multi-source data aggregation strategy (textbooks, web, public datasets) with source-specific preprocessing is thorough. The quality-aware stratification hierarchy (WeChat > Books > Paper-Single > Public Cls > Paper-Multi > Web) informed by human evaluation (Section 2.1) is a principled approach to managing heterogeneous data quality.

    3) strong empirical results with parameter efficiency: The performance gap between IRIS-4B and substantially larger models (e.g., +19.26 over Lingshu-32B, +23.65 over Qwen3-VL-32B in Table 1) is striking. Even IRIS-2B (overall avg. 69.57) outperforms all 32B/34B baselines. The consistent trend across model scales (2B/4B/8B) strengthens the claim that structured data quality matters more than parameter count for this domain.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1) Fundamentally unfail experimental comparision-closed-loop evaluaton on the authors’ own dataset: All baselines are evaluated in a zero-shot or pretrained setting on the IRIS-120K test set, while IRIS models are specifically fine-tuned on the IRIS-120K training set. The training and test data share the same data engine, source distribution, question formats, and generation templates. This creates a massive distribution advantage for IRIS that has nothing to do with clinical capability. A fair comparison would require either: (a) fine-tuning baselines (at least the same-architecture Qwen3-VL-4B) on the same training data, or (b) evaluating IRIS on external, independently curated OSD benchmarks. Without this, the claimed “SOTA performance” and “cross-scale suppression” are artifacts of distribution matching, not demonstrated clinical superiority. This alone substantially undermines the paper’s central claims.

    2) No external validation: The paper evaluates exclusively on IRIS-120K test data. There is no evaluation on any existing OSD benchmark, public ophthalmology dataset, or independent clinical test set. For a paper claiming “expert-level AI deployment” and “scalable OSD screening,” the absence of any external or clinical validation is a serious gap.

    3) Inappropriate metrics for clinical VQA tasks The open-ended evaluation relies on BLEU-1 and ROUGE-1/L-F (Table 1), which measure surface-level n-gram overlap. These metrics are poor proxies for clinical correctness—a response can be factually wrong while achieving high lexical overlap with a reference, and a clinically correct response may use different phrasing and score poorly. For a system intended for clinical deployment, domain-appropriate evaluation (e.g., clinical accuracy judged by ophthalmologists, factual correctness scoring, or at minimum GPT-based clinical correctness evaluation) is essential.

  • Please rate the clarity and organization of this paper

    Satisfactory

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The authors claimed to release the source code and/or dataset upon acceptance of the submission.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper tackles a genuinely important and underserved clinical problem, and the data curation effort behind IRIS-120K is substantial and well-designed. The Topic Finding Tree is a creative approach to injecting structured clinical reasoning into VQA data generation, and the ablation study provides useful evidence for its value.

    However, the experimental evaluation has fundamental validity issues that prevent acceptance in its current form. The central claim of “state-of-the-art performance” rests entirely on a closed-loop evaluation where IRIS is fine-tuned on IRIS-120K training data and evaluated on IRIS-120K test data, while all baselines are tested zero-shot. This comparison is inherently unfair and does not demonstrate that IRIS has learned genuine clinical reasoning rather than simply matching the data distribution of its own training set. The absence of any external validation, clinical evaluation, or fair fine-tuning comparison for baselines means the paper’s core claims are insufficiently supported.

    Additional concerns include: reliance on surface-level metrics (BLEU/ROUGE) for clinical tasks where correctness matters more than lexical overlap; no statistical analysis or uncertainty quantification; no limitations discussion; and unaddressed data provenance/ethics questions around web-scraped clinical images.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Reject

  • [Post rebuttal] Please justify your final decision from above.

    I maintain my score of 3.The rebuttal is candid and well-organized but confirms rather than resolves some core issues identified by the meta-reviewer and shared across all reviewers.

    The rebuttal discloses ophthalmologist review of a random sample (87/100 factual accuracy for captions, 4.4/5.0 for QA pairs), which partially addresses data quality concerns. However, this validates the generated dataset, not the model’s clinical outputs. Moreover, A single reviewer on a random sample does not constitute the systematic expert validation. This remains the key limitation of this work.

    My other concern is with the evaluation: The authors agree BLEU/ROUGE cannot measure diagnostic correctness and justify their use by field convention. The closed-ended tasks are stronger in principle but are themselves compromised by the circularity issue, near-perfect scores (97–99%) on a test set generated by the same data engine as the training set. No evaluation of model output clinical correctness (e.g., expert grading) is provided.

    The paper addresses a genuinely important and underserved clinical problem, and the data curation effort is substantial. The paper sits at the 3–4 boundary; I would not strongly object to acceptance if the camera-ready version substantially tempers the evaluation claims and clarifies the benchmark-specific nature of the results.



Author Feedback

We thank the Reviewers for recognizing the importance of external-eye OSD understanding, IRIS-120K, and the Topic Finding Tree/scene-driven design. We appreciate the concerns on dataset validity, comparison fairness, and clinical evaluation, and clarify them below.

Q1 Hallucinations in LLM/VLM-generated data@R1/R2/R4.We agree this is a key risk for any synthetic medical VQA corpus. In practice, our data were not generated from images alone. Each sample was anchored by original source information, including captions, disease labels, symptom descriptions, or textbook/public-dataset annotations; VLM descriptions were only auxiliary semantic surrogates. To reduce hallucination, we compared 5 captioning models and selected Qwen3-VL-32B considering quality and cost. A random subset reviewed by an ophthalmologist achieved 87/100 in factual accuracy. For TFT/QA generation, we similarly compared multiple LLMs and selected GPT-5-mini; ophthalmologist review of sampled QA pairs yielded 4.4/5.0.These details were omitted due to space. We acknowledge that “gold-standard” denotes a strictly filtered and deduplicated test set, not a fully multi-reader-adjudicated clinical benchmark. Still, source-quality stratification, anatomical TFT paths, pHash deduplication, and quality filtering jointly reduce residual noise. This follows common practices in recent medical VLM works such as LLaVA-Med, HuatuoGPT-Vision, and Lingshu, where GPT-series models are also used for data construction.

Q2 Closed-loop evaluation and fairness@R2/R4.We agree that our results demonstrate the value of IRIS-120K and structured instruction tuning, rather than universal superiority of the backbone. The manuscript includes the direct pretrained counterpart Qwen3-VL-4B under the same protocol; the gap between Qwen3-VL-4B and IRIS-4B isolates the effect of domain-specific supervision within the same backbone family. R2 suggests the near-perfect closed-ended scores may mainly reflect train/test distributional overlap. We clarify that large gains after domain-specific fine-tuning are common in specialized medical VLMs; for example, DentVLM reports nearly 100% improvement over the unfine-tuned Qwen3-7B in Fig. 3(b–d). Thus, strong gains from adapting a general VLM to a highly specialized visual domain are expected and do not by themselves imply leakage. We further used image-level deduplication and heterogeneous sources to avoid a narrow distribution of the test set. Public external OSD datasets are scarce (< 2000 images), so we merged them into IRIS-120K for greater diversity instead of relying on absent external validation. Fine-tuning every baseline would instead evaluate architecture-level differences under identical data and is infeasible for many closed or very large models. Thus, “SOTA” and “cross-scale” claims should be read as performance on the IRIS benchmark under this setting, not definitive superiority under all possible fine-tuning protocols, and external clinical validation remains important future work.

Q3 BLEU/ROUGE metrics@R1/R2/R4.We agree that n-gram overlap cannot fully measure diagnostic correctness, safety, or usefulness. We used these metrics due to reproducible and widely adopted in studies, including Hulu-Med, Lingshu, and MedDr. Importantly, our evidence is not solely based on them. Closed-ended questions evaluate anatomy- and pathology-specific decisions; TFT/Scene ablations show complementary contributions of structured reasoning and role adaptation; and expert sampling provides an additional clinical sanity check. Therefore, open-ended scores measure response alignment with references, not complete proof of clinical reliability.

In summary, IRIS is a data-centric contribution for an underserved ophthalmic modality. Its main claim is that clinically structured, anatomy-aware OSD instruction data can substantially improve a compact VLM over off-the-shelf general and medical VLMs. We thank the reviewers for their constructive feedback.




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    The paper addresses an important and underserved clinical problem: understanding ocular surface disease from external eye photographs, a modality underrepresented in current ophthalmic foundation models. Reviewers appreciated the clinical motivation, the scale of the IRIS-120K data effort, and the structured data-generation design using Topic Finding Trees and scene-driven interactions. The reported parameter efficiency of IRIS-4B is also promising.

    However, reviewers raised significant concerns about the validity and interpretation of the evidence. A central issue is that much of the dataset, including descriptions, reasoning chains, and VQA supervision, is generated and filtered by VLM/LLM-based pipelines, with limited evidence of systematic ophthalmologist verification, inter-rater agreement, or error analysis. Please clarify the extent of expert validation, especially for the final test set, and how clinically incorrect but plausible generated content is detected.

    Reviewers also questioned the fairness of the experimental comparison. IRIS is fine-tuned on the IRIS-120K training split and evaluated on a test split generated by the same data engine, whereas many baselines appear to be evaluated either zero-shot or pretrained. This makes it difficult to determine whether the reported gains reflect genuine clinical reasoning or distributional alignment with the generated dataset. Please clarify the baseline evaluation protocol and temper claims of SOTA/cross-scale superiority if the comparison is not controlled by comparable fine-tuning or external validation.

    Finally, reviewers noted that open-ended evaluation relies mainly on BLEU/ROUGE-style metrics, which are weak proxies for clinical correctness. Please address how the current evaluation supports claims about clinical reliability, interpretability, and deployment readiness. The rebuttal should focus on dataset/test-set validity, fairness of comparisons, and clinically meaningful evaluation.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    I recommend acceptance, subject to tempering the evaluation claims in the camera-ready version. Two of the three reviewers moved to Accept after the rebuttal, and the dissenting reviewer (R4) stated he would not strongly object to acceptance if the camera-ready substantially tempers the evaluation claims and clarifies that the results are benchmark-specific. The rebuttal addressed concerns by clarifying that the data are anchored to source captions, labels, and annotations rather than generated from images alone, with an ophthalmologist spot-check giving 87/100 caption factual accuracy and 4.4/5.0 for QA pairs, by including the same-backbone Qwen3-VL-4B under the same protocol to isolate the effect of domain-specific supervision, and by agreeing that the “SOTA” and “cross-scale” claims should be read as benchmark-specific rather than universal. The main contribution, a large structured instruction dataset for an underserved ocular-surface modality and evidence that such data improves a compact VLM, is valuable to the community. I therefore recommend acceptance with the explicit condition that the camera-ready clearly temper the SOTA and cross-scale claims to the IRIS benchmark setting, state the closed-loop nature of the evaluation as a limitation, and frame external clinical validation as necessary future work.



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    The paper is accepted as a borderline contribution. Two reviewers moved to or maintained acceptance after rebuttal, recognizing the importance of the ocular surface disease application and the value of the specialized dataset and lightweight model. One reviewer maintained rejection due to limited systematic expert validation, reliance on generated supervision, and potential circularity in the benchmark evaluation. However, the paper sits near the 3–4 boundary, and its clinical motivation and data curation effort are meaningful enough for acceptance, provided the final version clearly tempers claims about clinical reliability, emphasizes the benchmark-specific nature of the results, and acknowledges the need for stronger expert validation.



Meta-review #3

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    After considering the reviews and the rebuttal, I recommend acceptance. The paper addresses an important and underserved clinical problem, and the IRIS-120K data curation effort is substantial. The Topic Finding Tree and scene-driven design provide a useful way to structure the VQA data. The rebuttal clarifies the data generation process, expert sampling, comparison setting, and the intended scope of the results. I therefore recommend acceptance.



back to top