Abstract
Deep learning has become the dominant paradigm for medi- cal image analysis, yet its performance strongly depends on the availabil- ity of expert annotations. In dental imaging, where expert labels remain scarce and costly to obtain, several representation learning strategies have emerged to improve data efficiency, including supervised transfer learning, vision foundation models, and domain-specific self-supervised learning (SSL). However, their relative effectiveness for intra-oral pho- tographs remains poorly understood. In this work, we compare four representation learning strategies for intra-oral image analysis: training from scratch, ImageNet pretraining, DINOv2, and domain-specific Masked Autoencoder (MAE) pretraining. The comparison is performed across multiple annotation budgets on a private periodontitis screening dataset and validated on a public gingivi- tis benchmark. Our results show that general-purpose vision foundation models provide highly competitive representations for intra-oral photographs. Frozen DI- NOv2 generally outperformed the evaluated domain-specific MAE con- figurations, while the best overall performance was achieved by partially fine-tuning the last transformer blocks of DINOv2. Increasing the amount of unlabeled dental images used for MAE pretraining did not consistently improve downstream performance, and the evaluated MAE setups re- mained below the general-purpose pretrained baselines. These findings provide practical guidance for selecting representation learning strategies for annotation-efficient dental image analysis and high- light the potential of partially adapting large vision foundation models to intra-oral imaging.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026-sat/paper/ODIN_004.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to Open Review
Open Review Page: https://openreview.net/forum?id=vxrrc0BDQ3
BibTex
@InProceedings{MarNic_Comparing_MICCAISAT2026,
author = { Martin, Nicolas AND Chevallet, Jean-Pierre AND Mulhem, Philippe},
title = { { Comparing Pretraining Strategies for Intra-Oral Image Analysis Under Limited Annotation Budgets } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026 Workshops and Challenges},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 17270},
month = {pending},
page = {pending}
}
