Abstract

Automatic orthodontic report generation requires the joint interpretation of standardized intraoral photographs and three-dimensional intraoral scans (IOS). However, most multimodal large language models (MLLMs) are designed for two-dimensional visual inputs, making it dif- ficult to incorporate IOS geometry without a dedicated 3D encoder. We present an anatomy-aware rendering and visual fusion framework that converts maxillary and mandibular IOS meshes into clinically meaningful 2D representations. The framework establishes a standardized occlusal coordinate system, renders twelve views targeting whole-arch morphol- ogy and local occlusal relationships, and focuses the combined bite views on the dentition before integrating the renderings with intraoral pho- tographs. This design enables an existing dental-specialized MLLM to utilize complementary 3D geometric and 2D appearance information. Under the official Bite2Text evaluation against photograph-variant ref- erences, occlusal-region focusing raises the final score of the controlled twelve-view system from 0.3004 to 0.3363. Continuing fine-tuning from this focused model with photograph-report targets and applying topic- gated recall completion further raises the final score to 0.3811, outper- forming the zero-shot GPT-4o, GPT-5.5, and Qwen3-VL-8B baselines. These results demonstrate that clinically guided 3D-to-2D rendering and reference-target alignment provide an effective and practical interface be- tween IOS data and existing dental MLLMs.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/ODIN_challenges_007.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to Open Review

Open Review Page: https://openreview.net/forum?id=X60xeIP9KX

BibTex

@InProceedings{GuoJia_Clinically_MICCAISAT2026,
        author = { Guo, Jiashuo AND Hao, Jing AND Hung, Kuo Feng AND Wang, Feng},
        title = { { Clinically Guided Rendering for Orthodontic Report Generation } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 17270},
        month = {pending},
        page = {pending}
}


back to top