List of Papers Browse by Subject Areas Author List
Abstract
Accurate and clinically relevant non-invasive aging assessment remains a key challenge in gerontology. While computer vision enables functional phenotypic analysis, existing methods primarily rely on unimodal analysis, struggle to encapsulate the systematic and complex nature of the aging process. We propose a hierarchical cross-modal fusion framework that integrates three complementary aging-related modalities: optical coherence tomography angiography (OCTA) and optical corneal thickness (OCT) from the ocular modality, millisecond-level behavioral syllables from the behavioral modality, and a high-dimensional gait matrix from the gait modality. Our framework introduces specialized training paradigms for each modality, including a three-stage progressive self-masked training paradigm, an active learning behavioral representation framework, and a multi-stage gait quantization workflow. This fusion architecture systematically progresses from unimodal encoding and intra-domain alignment of OCTA-OCT to inter-domain feature enhancement, and employs an entropy-guided probabilistic weighting strategy to integrate peripheral biological structural changes with central functional outputs. We performed aging assessments using a novel dataset of 166 mice aged 8 weeks to 18 months, divided into four age groups, with concurrent ocular, behavioral, and gait measurements. Experimental results demonstrate that our method achieves an accuracy (ACC) of 92.31\% and an area under the curve (AUC) of 97.73\% in aging assessment, significantly outperforming both single-modal baseline methods and state-of-the-art approaches. Furthermore, it exhibits superior capability in the precise extraction of age-related features. This work establishes a clinically interpretable framework for non-invasive aging assessment by bridging structural and functional changes at different biological scales.
Code is available at \url{https://github.com/hangxingwu/HCMFMSTP}.
Links to Paper and Supplementary Materials
Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/2578_paper.pdf
SharedIt Link: Not yet available
SpringerLink (DOI): Not yet available
Supplementary Material: Not Submitted
Link to the Code Repository
https://github.com/hangxingwu/HCMFMSTP
Link to the Dataset(s)
N/A
BibTex
@InProceedings{WuYih_Hierarchical_MICCAI2026,
author = { Wu, Yihang AND Xu, Junnan AND Li, Guocong AND Zhang, Yiwen AND Cao, Bo AND Huang, Shuaihan},
title = { { Hierarchical Cross-Modal Fusion with Modality-Specific Training Paradigms for High-Precision Non-invasive Aging Assessment } },
booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
year = {2026},
publisher = {Springer Nature Switzerland},
volume = {LNCS 16896},
month = {September},
page = {pending}
}
Reviews
Review #1
- Please describe the contribution of the paper
This paper proposes a noninvasive aging assessment method for mice based on four modalities, including imaging and behavioral data such as OCT, OCTA, behavioral syllables, and gait matrices.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
1.Proposed a noninvasive aging assement method using cross-modal fusion methods. 2.Applied self-masking pre-training method. 3.Applied cross-modal fusion method for multiple input modalities and two other modalities. 4.This method can be extended to humans for aging assessment using OCT and OCTA imaging modalities to evaluate retinal health and predict age for datasets lacking age information.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
The paper exhibits several significant weaknesses that limit its reliability and clarity. The dataset is very small (166 mice), resulting in an extremely limited number of test samples per class, which undermines the validity of the reported performance; n-fold cross-validation is needed to better assess robustness. Additionally, the manuscript lacks critical methodological details, including the definition and implementation of key components (e. g. , hierarchical cross-modal fusion, weighting function (w(i)), and probability estimation), as well as insufficient explanation of model variants and inputs (e. g. , OCT/OCTA usage).
More detailed comments are as follows: 1.The dataset is too small, consisting of only 166 mice. Since 20% of the data is used for testing and there are four classes, each class has only about 8 samples in the test set. This is insufficient to reliably evaluate the performance of the proposed method. N-fold cross-validation should be performed to better demonstrate its effectiveness. 2.The classification performance for each individual modality should be reported to enable a more detailed analysis of the model and its features. 3.”Hierarchical Corss-Model Fusion” did not provide detail information. It is difficult to understand the benefit of using the aggregation module (M) output as input to the MLPs for OCT and OCTA, since both MLPs are trained on individual modalities. The authors should clarify and justify the advantage of incorporating the aggregated results. 4.Please specify how 𝑤(𝑖) is applied in the process, indicate the value of 𝑇 used in the experiments, and specify how to estimate P(i). Where is this process in the proposed workflow? 5.Please clarify the meaning of the hierarchical relationships between modalities used in the proposed method. I could not find a hierarchical structure in the proposed method. 6.Please specify the number of OCT and OCTA images used per mouse as input. 7.OCT was mis-spelled ad “CT”. 8.There is no explanation of the differences between Ours-B, Ours-L, and Ours-H. How were these three variants implemented with different parameters? 9.Is it true that all metrics have the same values for Ours-B, Ours-L, and Ours-H? 10.In Figure 3, which variant of the Ours model did the authors use to generate the heatmaps?
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not provide sufficient information for reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(5) Accept — should be accepted, independent of rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
I would like to recommend acceptance of this paper, as the proposed method presents a non-invasive cross-modal fusion approach for age assessment. Such a non-invasive framework has strong potential for safe and effective application to human subjects.
The work is particularly interesting in that combining imaging data with gait and behavioral information significantly improves performance, even though the gait and behavioral modalities alone show relatively limited predictive power. This highlights the value of cross-modal integration and suggests meaningful complementary information across modalities.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
Accept
- [Post rebuttal] Please justify your final decision from above.
The proposed non-invasive framework has strong potential for safe and effective application in human subjects.
Review #2
- Please describe the contribution of the paper
This study proposed a deep learning fusion model that integrated multimodal data, including ocular features (OCT and OCTA), behavioral features, and gait features. The study was conducted using a mouse model, and the measurements appeared to be reliable. The topic of estimating aging using features from multiple domains was important and timely. The experimental design was reasonable, and the results demonstrated the advantages of the proposed model.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The major strength of this study was the multimodal feature fusion design. Overall, the authors provided detailed descriptions of each modality, and the model design appeared to be reasonable. Prior studies have primarily focused on using behavioral features or ocular features alone to estimate aging, whereas this study incorporated ocular, behavioral, and gait data in a unified framework. The selection of target age groups appeared to be appropriate. The dataset split into training, validation, and test sets using a 7:1:2 ratio was also reasonable. The quantitative evaluations appeared to be appropriate. Overall, the study followed standard methodological practices.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
I do not see any major weaknesses in this study. There are only a few places where clarity could be improved. 1) The description of how the three modalities were fused could be clearer. Figure 1d shows a normalized sum, while Section 2.3 describes hierarchical cross-modal fusion using feature distribution entropy, and the relationship between these two descriptions was not entirely clear. 2) In Table 1, the distinctions among “Ours-B,” “Ours-L,” and “Ours-H” were unclear. Although these variants appeared to differ in the number of parameters, they showed identical performance across all reported metrics, which would benefit from further explanation. 3) In Table 2, it would be helpful to include additional combinations such as “Ocular + Behavior” and “Ocular + Gait” to provide a more complete assessment of modality contributions. In addition, using “Ocular” instead of “Ours” for labeling would improve clarity and consistency. 4) In Section 2.1, it was not clear how the ROI was obtained. 5) There was also a minor typo in Section 2.1, where “CT” should have been “OCT.”
- Please rate the clarity and organization of this paper
Good
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not provide sufficient information for reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(5) Accept — should be accepted, independent of rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
This study presented a multimodal fusion framework that integrated ocular (OCT and OCTA), behavioral, and gait features, which extended prior work that often focused on a single modality. The methodology was generally well described, and the model design and experimental setup appeared to be reasonable. The use of a mouse model and the reported measurements suggested acceptable reliability. In addition, the dataset partitioning and evaluation strategy followed standard practices, and the results indicated potential advantages of the proposed approach. Overall, the topic of aging estimation using multimodal data was relevant and of current interest.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
Accept
- [Post rebuttal] Please justify your final decision from above.
I thank the authors for providing additional information to address the reviewers’ comments. Overall, I understand that the limited data size does constrain the generalizability of this approach, but the hierarchical cross-modal fusion also shows strong potential for future work. The authors also stated that they will release their code. I believe this paper should be accepted.
Review #3
- Please describe the contribution of the paper
This paper proposes a hierarchical cross-modal fusion framework for non-invasive aging assessment by integrating ocular (OCT/OCTA), behavioral, and gait modalities. The method introduces modality-specific training paradigms and an entropy-guided fusion strategy. Experiments on a self-collected dataset of 166 mice show strong performance improvements over unimodal and existing methods.
- Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
The paper presents a multimodal aging assessment framework by integrating ocular, behavioral, and gait information, providing a more comprehensive characterization of aging. It introduces modality-specific training strategies tailored to different data types and proposes a staged cross-modal fusion pipeline. Additionally, an entropy-based dynamic weighting mechanism is used for multimodal fusion. The construction of a multimodal dataset also contributes to this research direction.
- Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.
1.The authors constructed a multimodal dataset integrating ocular, behavioral, and gait information, which is valuable. However, the overall dataset size is relatively limited, which may restrict the robustness and generalization ability of the proposed model. 2.There are minor writing issues in the manuscript. For example, “CT data” should be corrected to “OCT data” to ensure consistency and accuracy. 3.The method relies on a pre-training and fine-tuning paradigm, which is a widely adopted strategy. However, the pre-training dataset is not publicly available, and the framework (Fig. 1) appears to use natural images for pre-training. Such a setting may introduce domain gaps, and simple fine-tuning may not be sufficient for ophthalmic or medical imaging tasks. It would be more appropriate to leverage domain-specific pre-training, such as ophthalmic foundation models or medical vision-language models, which may further improve performance. 4.The claimed “Hierarchical Cross-Modal Fusion” is not clearly reflected in the model design. The current entropy-based fusion strategy appears relatively simple and may be insufficient to capture complex relationships between structural degeneration and functional decline. Moreover, the benefits of cross-modal fusion are not convincingly demonstrated. The ablation study lacks comparisons with more advanced fusion strategies (e.g., attention-based or transformer-based fusion methods). 5.The experimental setup lacks rigorous validation. Although a train/validation/test split is used, multi-fold cross-validation is not conducted, which may lead to performance variability. In addition, the model is heavily evaluated on a private dataset without external validation, raising concerns about its generalization capability. 6.The paper lacks modality contribution analysis. It remains unclear which modality plays a dominant role in the final prediction, and how each modality contributes to the overall performance. Such analysis would improve the interpretability and credibility of the proposed framework.
- Please rate the clarity and organization of this paper
Satisfactory
- Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.
The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.
- Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?
N/A
- Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html
N/A
- Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.
(3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal
- Please justify your recommendation. What were the major factors that led you to your overall score for this paper?
Although a novel dataset is constructed, the methodological innovation is limited, and the comparisons of multimodal fusion strategies are insufficient. The experiments rely on a relatively small private dataset, lacking external validation and multi-fold cross-validation, which raises concerns about generalizability and robustness. In addition, the work lacks in-depth analysis of clinically meaningful findings.
- Reviewer confidence
Very confident (4)
- [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.
Accept
- [Post rebuttal] Please justify your final decision from above.
N/A
Author Feedback
Q1: Small dataset and lack of external testing (MR1, R1, R3). A: First, collecting mouse datasets is highly challenging, and public datasets are difficult to obtain. Our dataset is the first cross-modal dataset involving ocular, behavioral, and gait modalities in mice. It includes approximately 200 mice, collected over one year in three batches. For each mouse, OCT/OCTA images were acquired five times for both the left and right eyes for manual screening, resulting in 324 high-quality OCT/OCTA images. Second, mice have an approximately 20% mortality rate during anesthesia and breeding; therefore, we only counted the 166 surviving mice. In addition, mouse-methodology and clinically related studies generally include 5-50 mice, so the scale of our dataset is already relatively large and covers the entire lifespan of mice. Finally, under the small-sample setting, we evaluated model robustness through cross-batch validation and 5/3-fold cross-validation experiments. Q2: Details and advantages of Hierarchical Cross-modal Fusion(HCF) (MR1, R1, R2, R3).A: Unlike previous single-level fusion methods, HCF contains three levels of fusion: 1) intra-subject fusion, where feature maps of the two corresponding left/right-eye images within the same modality are normalized and aggregated to reduce the influence of acquisition noise and binocular differences; 2) intra-ocular-domain fusion, where OCTA and OCT are aggregated by M at the flattened vector level into a shared ophthalmic-domain representation, thereby integrating complementary aging features from structural and vascular information; 3) cross-domain fusion, where the entropy values of the probability distributions of the four modalities are computed separately, and the softmax-normalized weights w(i)(T=1 to ensure the true distribution) are used as the weights for the probability distribution of each age group C under each modality, so that the probability that the mouse age is c can be obtained by P_c = Σ_i w_i*p_{i,c}, where i ∈ {1, 2, 3, 4}. Notably, the MLPs ensure that the complementary information of the two modalities can be fully exploited. Since the OCT and OCTA encoders have already completed joint self-masked fine-tuning in the second stage, their output embeddings lie in the same shared subspace. Therefore, the two MLP layers in the third stage also jointly account for the processing of corneal and vascular information, although with different emphases. Q3: Distinction and analysis of HCF variants (MR1, R1, R2). A: The three variants respectively use ocular-modality encoders identical to ViT-Base/Large/Huge, while the decoder is fixed to 8 Transformer decoder blocks and a channel width of 512.All parameters were randomly initialized. Their comparable performance indicates that HCF is the main factor contributing to the performance improvement. Q4: Experimental details (MR1, R1, R2, R3). A: For the gait and behavior modalities, we selected more than 10 machine-learning models for 5-fold cross-validation experiments and reported the average cross-validation results of the best-performing model. For the ocular modality, we performed 3-fold cross-validation for each model and reported the average results. The heatmaps in Fig. 3 were generated by Ours-B. Ours-L and Ours-H also show similar effects, and even outperform directly fine-tuned ophthalmic foundation models such as RETFound/VisionFM. Due to space limitations, we did not report the ablation results for more modalities. Specifically, the performance of “Gait” and “Behavior” decreased to varying degrees compared with “Gait+Behavior”; the performance of “Ocular+Behavior” and “Ocular+Gait” was between that of “Ocular” and “Full Method”. Q5: Generalizability and limitations (MR1, R1, R3). A: HCF has already been applied to the preclinical drug screening stage for metformin and also shows strong promise for clinical-stage applications. We will release our code upon acceptance and revise the manuscript accordingly.
Meta-Review
Meta-review #1
- Your recommendation
Provisional Accept
- Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.
The paper proposes a multimodal framework for non-invasive aging assessment by integrating ocular, behavioral, and gait data. Reviewers agree that the topic is relevant and that multimodal fusion is a promising direction, supported by an interesting dataset.
However, several important concerns are raised. The dataset is relatively small (166 mice), raising significant questions about robustness and generalization. The experimental protocol lacks stronger validation, such as cross-validation and external testing. In addition, the methodological contribution is not fully clear: the “hierarchical cross-modal fusion” is insufficiently described and not convincingly demonstrated beyond relatively simple fusion strategies. There is also a lack of modality contribution analysis and unclear distinctions between model variants.
The authors should address the following points: 1- Clarify the proposed fusion method and its advantage over simpler or existing approaches. 2- Provide clearer analysis of modality contributions and model variants. 3- Improve reproducibility and clarify missing methodological details. 4- Discuss generalization and limitations more explicitly.
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Accept
- Please justify your recommendation.
The rebuttal addressed the main concerns reasonably well. In particular, the authors clarified the structure of the hierarchical cross-modal fusion, explained the meaning of the weighting scheme and the role of the different fusion levels, distinguished the model variants, and provided additional details on validation protocols and modality contributions. They also clarified that robustness was assessed with cross-batch validation and cross-validation, which helps mitigate concerns about the relatively small dataset.
The dataset size and lack of external validation remain limitations, and the methodological novelty is still moderate. However, these concerns are outweighed by the novelty of the multimodal dataset, the practical interest of the application, and the satisfactory clarifications in the rebuttal.
Meta-review #2
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Accept
- Please justify your recommendation.
This paper makes a meaningful contribution by integrating ocular, behavioral, and gait information for non-invasive aging assessment in a unified multimodal framework, and the reviewers consistently recognized the application potential. The rebuttal appears to have helped by clarifying several design details and reinforcing the practical motivation of the work, and the post-rebuttal reviewer opinions remained supportive. Although limitations around scale and broader validation remain, it is a good paper with contribution to multimodal data analysis.
Meta-review #3
- After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.
Accept
- Please justify your recommendation.
The paper addresses an important problem (non‑invasive aging assessment) with a clinically meaningful multimodal framework integrating ocular, behavioral, and gait data. The dataset, while moderately sized (166 mice), represents a substantial collection effort covering the full lifespan, and the hierarchical cross‑modal fusion design is well motivated. Both Reviewer #1 and Reviewer #2 gave Strong Accept (5), recognizing the practical value and solid empirical results. Reviewer #3’s Weak Reject (3) raised concerns about dataset size and external validation, but these limitations do not invalidate the core contribution, and the two Accept scores provide a clear positive consensus. Therefore, the paper is accepted based on the original submission, without relying on new results introduced during rebuttal.
