Abstract

Smartphone-based screening for Autism Spectrum Disorder (ASD) provides a scalable and low-cost alternative to traditional clinical assessments. However, unconstrained mobile data are often noisy and exhibit high inter-subject behavioral variability, which undermines the effectiveness of previous unimodal and static fusion methods. To address these challenges, we propose an adaptive multimodal fusion network that jointly models eye-movement patterns and temporal head-pose dynamics to capture instance-wise complementary behavioral cues. Specifically, we first extract high-level eye-movement features using a dual-branch network that encodes both static fixation and dynamic scanpath information. We then introduce a joint representation gating module for instance-level adaptive fusion, which dynamically reweights modalities to suppress unreliable signals and enhance discriminative cues. Additionally, we incorporate supervised contrastive learning as an auxiliary objective to structure the latent space, thereby improving intra-class compactness and inter-class separability, as well as overall generalization in limited and heterogeneous medical datasets. Experiments on our collected real-world mobile dataset of 124 children demonstrate state-of-the-art performance, achieving 91.93% accuracy and an F1-score of 0.9206.

Links to Paper and Supplementary Materials

Main Paper (Open Access Version): https://papers.miccai.org/miccai-2026/paper/2449_paper.pdf

SharedIt Link: Not yet available

SpringerLink (DOI): Not yet available

Supplementary Material: Not Submitted

Link to the Code Repository

https://github.com/JIESunshine/AMF-Net

Link to the Dataset(s)

N/A

BibTex

@InProceedings{XuJie_SmartphoneBased_MICCAI2026,
        author = { Xu, Jie AND Guo, Xinran AND Tang, Longbin AND Li, Kuan AND Xia, Chen},
        title = { { Smartphone-Based ASD Screening via Instance-Adaptive Cross-Modal Fusion and Contrastive Learning } },
        booktitle = {Medical Image Computing and Computer Assisted Intervention -- MICCAI 2026},
        year = {2026},
        publisher = {Springer Nature Switzerland},
        volume = {LNCS 16896},
        month = {September},
        page = {pending}
}


Reviews

Review #1

  • Please describe the contribution of the paper

    The authors propose an end-to-end multimodal framework that integrates eye-movement and head-pose cues for ASD screening (binary classification) in unconstrained mobile scenarios. In a novel applicative fashion, if compared with recent the state of the art, the authors also introduce a customized instance-level adaptive cross-modal fusion via the Joint Representation Gating module to dynamically adjust modality contributions, and incorporate cross-modal contrastive learning to enhance feature representation. Potentially, the contributions may lie in both application and methodological sides.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.
    • The clinical task is well motivated and introduced within the recent state of the art.
    • The methodological challenges to be addressed with respect the state of the art are very clear.
    • At high level, the methodological framework is properly described in each component, as well as the incremental novel technical contributions, properly highlighted by the authors with respect of the recent state of the art for the specific clinical task.
    • Potential real-world clinical impact
  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    Major points:

    • Code not shared by authors.
    • Lack of ethics considerations for experimental procedures where human subjects are involved in.
    • Lack of ethics protocol approval statement.

    Minors point:

    • Lack of feedback/references for a possible deployment phase.
    • Lack of standard deviations in 5-fold cross validation experimental results.
    • Lack of statistical analysis in the experimental results to claim, if exist, statistically significant differences.
    • Lack or partial missing information about hyperparameters optimization, experimental metric to be maximized, definition of micro/macro metrics, etc.
  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not provide sufficient information for reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    Considering the involvement of human subject data, major information regarding subject consent and ethical approval should be provided.

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    NA

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (3) Weak Reject — marginally below the acceptance threshold, but would not mind if accepted, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    Even if methodological procedure is well motivated and step-by-step explained at high-level, the authors do not release the code, which would at least allow for verification of the methodological procedure and some technical choices that are not fully presented or clearly explained in the paper. Given the high-quality standards of this conference and, first of all, the principles of transparent and reproducible research, the reviewer believe this is the most significant drawback of this work, coupled with absence of ethics considerations due to experiments where human subjects are involved. Given the previous motivations, even if the clinical motivation is well motivated as well the applicative methodological improvement with respect the state of the art, the lack in both reproducibility and ethics considerations may suggest a weak reject decision.

  • Reviewer confidence

    Somewhat confident (2)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The authors have addressed my primary concerns, particularly regarding ethics compliance (they confirm that all procedures were approved by the institutional Medical Ethics Committee, followed the Declaration of Helsinki, and written informed consent was obtained from the legal guardians of all child participants). Regarding reproducibility, the authors commit to releasing complete code, preprocessing scripts, training configurations, and de-identified data upon acceptance, which resolves my main reservation. Moreover, the additional details provided (standard deviations to show cross-fold stability, statistical significance tests with p<0.05, details on comparisons) strengthen confidence in the methodology. Overall, the study presents an interesting and clinically relevant contribution for mobile ASD screening.



Review #2

  • Please describe the contribution of the paper

    This paper proposes an adaptive multimodal framework for Autism Spectrum Disorder (ASD) screening that jointly leverages eye movement patterns and temporal head pose dynamics to capture instance-specific behavioural cues. The framework incorporates a Joint Representation Gating (JRG) module that dynamically adjusts modality-specific contributions at the subject level, addressing the inherent challenges posed by behavioural heterogeneity and environmental noise in naturalistic ASD assessment settings. A supervised contrastive learning objective is employed to improve the quality and discriminability of the learned representation space.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    1.The framework addresses need of combining eye-movement and head-pose modalities for ASD screening. 2.The proposed JRG module is a notable contribution, enabling the model to dynamically and adaptively weight modality-specific contributions on a per-subject basis. This directly addresses the well-known challenge of behavioural heterogeneity in ASD populations, where the relative informativeness of eye-movement versus head-pose signals may vary substantially across individuals. 3.The ablation study in Table 2 is a strong addition, clearly quantifying the contribution of each component of the proposed method.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    1.While Section 2.1 describes the subject split, no ethical consent declaration is present in the manuscript. For studies involving human participants, particularly a vulnerable population such as children with ASD, this is a critical omission that must be addressed prior to publication. Additionally, the authors should clarify: (a) the clinical basis for ASD and TD categorization (e.g., ADOS scores, formal clinical diagnosis), (b) the specific video stimuli chosen for data collection and the rationale behind this choice, supported by relevant references, and (c) any inclusion/exclusion criteria applied during participant recruitment.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    1.Figure 3 is difficult to interpret due to its small size and the reliance on shape alone to distinguish data points or categories. Given that shape differentiation at reduced scale is perceptually demanding, the authors are strongly recommended to show categories using distinct colors alongside shape markers.

    2.Limited Reproducibility: While a full dataset release may not be feasible, releasing the processing and training code even without data would significantly increase the paper’s utility to the research community and is strongly encouraged.

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    The paper presents a well-motivated and technically sound framework for multimodal ASD screening, with meaningful contributions in the form of the JRG module and the use of supervised contrastive learning. The ablation study is thorough, and the methodological choices are clearly explained. However, the absence of an IRB statement and clinical categorization details is a significant concern for a study and must be rectified. The readability of Figure 3 is a presentation issue that should be straightforward to address.

  • Reviewer confidence

    Very confident (4)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    I thank the authors for their thorough response, which successfully addresses all of my concerns. The authors have clearly clarified the clinical protocols for ASD/TD grouping using ADOS-2, justified the stimulus design with robust literature, and provided inclusion/exclusion criteria. Furthermore, their commitment to improving Figure 3’s visual clarity and their commendable plan to openly release the complete source code, preprocessing scripts, and de-identified data significantly enhance the paper’s transparency and reproducibility. I recommend this paper for publication.



Review #3

  • Please describe the contribution of the paper

    This paper introduces a smartphone-based multimodal framework for screening Autism Spectrum Disorder (ASD) using eye-movement and head-pose cues extracted from unconstrained mobile recordings. It’s main contribution is an instance-adaptive fusion strategy built around a Joint Representation Gating (JRG) module, which assigns modality weights on a sample-by-sample basis rather than relying on simple feature concatenation or fixed global fusion weights. The framework is further enhanced by a supervised contrastive learning objective, designed to improve the alignment of multimodal representations while also increasing class separability. Evaluated on a real-world dataset of 124 children, the proposed method achieves stronger classification performance than several existing unimodal and multimodal baselines. Overall, the results suggest that adaptive fusion is a promising direction for ASD screening, particularly in settings characterized by behavioral heterogeneity and varying signal quality.

  • Please list the major strengths of the paper: you should highlight a novel formulation, an original way to use data, demonstration of clinical feasibility, a novel application, a particularly strong evaluation, or anything else that is a strong aspect of this work. Please provide details, for instance, if a method is novel, explain what aspect is novel and why this is interesting.

    The paper studies an important and clinically relevant problem. Smartphone-based ASD screening is clearly useful in practice because it could make screening more accessible and easier to deploy at scale than lab-based methods. I also think the paper frames the task well by focusing on unconstrained mobile data collection, where issues like noise, imperfect calibration, and large behavioral variation across children are unavoidable and should be treated as part of the problem.

    The multimodal setup makes sense for this application. Combining eye-movement and head-pose information is a reasonable design choice, especially since gaze estimates from mobile devices can be noisy and not every child will express ASD-related behaviors in the same way. The eye-movement branch is also well designed. Using fixation heatmaps and scanpath maps is more robust than feeding raw gaze coordinates directly into the model, particularly in a less controlled mobile setting.

    Joint Representation Gating module is a strong part of the paper. The main idea of letting the model decide how much to rely on each modality for each sample is well matched to the task, since the quality of gaze and head-pose signals can vary a lot from one recording to another. This is also supported by the ablation results, which show that the adaptive fusion strategy works better than simple concatenation, equal weighting, or a single set of learned fusion weights.

    The paper does a good job with ablations and provides some useful interpretability analysis. The modality, fusion, and loss ablations all support the main argument of the paper. The paper gives good insight into how the adaptive fusion mechanism behaves for ASD versus TD samples.

  • Please list the major weaknesses of the paper. Please provide details: for instance, if you state that a formulation, way of using data, demonstration of clinical feasibility, or application is not novel, then you must provide specific references to prior work.

    The main weakness is that the methodological novelty is limited. Both adaptive multimodal gating and supervised contrastive learning are already well established. The contribution is mostly in combining these ideas for smartphone-based ASD screening rather than introducing a genuinely new method. That is acceptable for an application-driven MICCAI paper, but the novelty claims should be stated more carefully.

    The baseline comparison could also be stronger. While several prior methods are included, it is not fully clear whether they were evaluated under the same preprocessing, splits, and training setup. More importantly, the paper does not compare against stronger generic fusion methods built on the same unimodal encoders, such as cross-attention, FiLM-style fusion, or uncertainty-aware fusion. As a result, it is difficult to determine how much of the gain comes specifically from the JRG module.

    The evidence for the adaptive fusion claim is also somewhat incomplete. The paper argues that JRG downweights noisy modalities and relies more on stronger signals, which is plausible, but this is not tested directly. A more convincing analysis would corrupt or remove one modality in a controlled way and check whether the learned weights respond as expected.

    Reproducibility is reasonable but not complete. The overall architecture is described, but some implementation details remain vague, including parts of the contrastive setup, handling of low-confidence gaze points, and parts of preprocessing.

  • Please rate the clarity and organization of this paper

    Good

  • Please comment on the reproducibility of the paper. Please be aware that providing code and data is a plus, but not a requirement for acceptance.

    The submission does not mention open access to source code or data, but provides a clear and detailed description of the algorithm to ensure reproducibility.

  • Based on your review and your understanding of the MICCAI Scientific Code of Ethics, do you believe this submission may involve a potential ethics concern or violation?

    N/A

  • Optional: If you have any additional comments to share with the authors, please provide them here. Please also refer to our Reviewer’s guide on what makes a good review and pay specific attention to the different assessment criteria for the different paper categories: https://conferences.miccai.org/2026/en/REVIEWER-GUIDELINES.html

    N/A

  • Rate the paper on a scale of 1-6, 6 being the strongest (6-4: accept; 3-1: reject). Please use the entire range of the distribution. Spreading the score helps create a distribution for decision-making.

    (4) Weak Accept — marginally above the acceptance threshold, but would not mind if rejected, dependent on rebuttal

  • Please justify your recommendation. What were the major factors that led you to your overall score for this paper?

    A weak accept because the paper addresses an important clinical problem and proposes a coherent multimodal solution that is well aligned with the challenges of unconstrained mobile data. The instance-adaptive fusion mechanism is practically motivated, the use of eye-movement and head-pose cues is sensible, and the experimental results suggest a meaningful improvement over the included baselines. The paper is also generally well organized and includes relevant ablations that support the main message.

  • Reviewer confidence

    Confident but not absolutely certain (3)

  • [Post rebuttal] After reading the authors’ rebuttal, please state your final opinion of the paper.

    Accept

  • [Post rebuttal] Please justify your final decision from above.

    The clarification that the contribution is positioned as an application-oriented adaptation of adaptive gating and supervised contrastive learning for mobile ASD screening, rather than a fundamentally new method, is reasonable, and the commitment to refine the contribution statements accordingly is appropriate. The additional comparisons against Cross-Attention, FiLM, and Uncertainty-Aware Fusion under the same backbone and protocol directly address my concern about baselines and provide more convincing evidence that the gains stem from the JRG module. The modality perturbation experiments, showing that progressive degradation of one modality systematically reduces its learned weight, also provide the controlled evidence I was looking for in support of the adaptive fusion claim. Together with the planned code and data release, these clarifications sufficiently address my earlier points, and I maintain my recommendation of weak accept (4).



Author Feedback

We sincerely thank all reviewers for their valuable time and comments. We are encouraged by the positive comments regarding novelty (R1&R2), practical applicability (R1&R3), and experiments/presentation (R2&R3).

Ethics Approval (R1&R2): All ethical and experimental procedures were approved by the institutional Medical Ethics Committee and followed the latest version of the Declaration of Helsinki. Written informed consent was obtained from the legal guardians of all child participants before participation. We will clarify the ethics approval and informed consent statement in the manuscript.

Reproducibility (R1-R3): We will release the complete code, preprocessing scripts, training configurations, and experimental stimulus videos upon acceptance. These materials will cover supervised contrastive learning, low-confidence gaze handling, gaze/scanpath-map generation, and head-pose preprocessing. De-identified processed eye-movement and head-pose data will also be provided as permitted by ethics approval and informed consent.

R1: References for Deployment: AMF-Net only requires a 4-minute video recorded by a mobile-device camera. We have provided supporting references for its key components, including mobile gaze estimation [27], head-pose extraction [1], and supervised contrastive regularization [13].

Standard Deviations: Five-fold standard deviations for Accuracy, Precision, Recall, and F1-score are 0.0438, 0.0682, 0.0679, and 0.0454, respectively, showing competitive cross-fold stability.

Significance Test/Settings/Metrics: AMF-Net significantly outperforms others under the paired t-test (p<0.05), with 4.97% and 4.06% improvements in Accuracy and F1-score while maintaining a more balanced Precision-Recall trade-off. The main hyperparameters and training settings are provided in Section 3.1, and we follow the same or equivalent metrics and protocols as prior ASD screening studies [4,7,22,26].

R2: ASD/TD grouping was determined using ADOS-2 assessments by qualified clinicians and clinical diagnostic information. The calibration stimuli followed prior gaze-estimation work [27], while the test stimulus concatenated social and geometric scenes (Section 2.1 & Fig. 1), motivated by abnormal preference for geometric stimuli as an ASD biomarker (Pierce et al., Biological Psychiatry, 2016). Inclusion criteria were age 2-8 years and normal/corrected vision; exclusion criteria included age outside this range, inability to cooperate, or visual problems such as strabismus. We will improve Fig. 3 by enlarging fonts/labels and using distinct colors to distinguish the two modalities.

R3: Novelty: The main novelty of this work lies in developing an application-oriented model for real-world mobile ASD screening. Based on modality-specific encoding, we adapt and integrate established adaptive multimodal gating and supervised contrastive learning to address practical challenges in unconstrained mobile scenarios, including noisy gaze estimation, large head-pose variations, and inter-subject behavioral heterogeneity. We will refine the contribution statements to better highlight the task-specific adaptation of these components for mobile ASD screening.

Fusion Comparison: All compared methods use the same preprocessing, subject-level five-fold split, and training/evaluation settings. We compare AMF-Net with Cross-Attention, FiLM, and Uncertainty-Aware Fusion under the same backbone. They achieve Accuracy/F1-score of 0.8873/0.8848, 0.8797/0.8703, and 0.8870/0.8894, respectively, while AMF-Net achieves 0.9193/0.9206, confirming the effectiveness of JRG.

Adaptive Fusion: ASD children have fewer valid eye-movement points than TD children, with Fig. 2 showing lower eye-movement weights for ASD cases. Moreover, modality perturbation experiments reveal that progressive degradation of one modality gradually reduces its weight while raising the other’s weight. These results confirm JRG adaptively adjusts modality contributions by signal reliability.




Meta-Review

Meta-review #1

  • Your recommendation

    Invite for Rebuttal

  • Please justify your decision. In case you deviate from the reviewers’ recommendations, explain in detail the reasons why. In case of an invitation for rebuttal, clarify which points are important to address in the rebuttal.

    This submission proposes an adaptive multimodal fusion network for smartphone-based ASD screening. Several critical concerns remain regarding ethics, dataset details, novelty, and reproducibility. The key points to address in the rebuttal are as follows: Both R1 and R2 highlight the absence of ethical considerations for experiments involving human subjects. R1 and R2 also request more details about the dataset. Both R1 and R3 raise concerns about reproducibility. R3 notes that the method mainly combines existing components.

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    All three reviewers recommend acceptance (after rebuttal). The responses addressed the concerns about ethics approval and reproducibility.



Meta-review #2

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    Post rebuttal, all three reviewers are leaning towards accepting this work. Major review concerns seem to have been sufficiently addressed- through questions on reproducibility and code sharing, ethical considerations involving human subjects still remain.



Meta-review #3

  • After you have reviewed the rebuttal and updated reviews, please provide your recommendation based on all reviews and the authors’ rebuttal.

    Accept

  • Please justify your recommendation.

    Based on the three post-rebuttal reviews, all reviewers are now supportive of acceptance. Reviewer #1 highlights satisfactory resolution of ethics and reproducibility concerns, endorsing acceptance. Reviewer #2 states that the authors thoroughly addressed all concerns, including clinical protocols, stimulus design, and data/code release, and recommends publication. Reviewer #3 acknowledges the reasonable positioning as an application-oriented adaptation, the added baseline comparisons, and the modality perturbation experiments, maintaining a weak accept recommendation.



back to top