作者
Wondimagegn Bekele Munto,M. Zhou,Yi Pan,Yanjie Wei,Wenhui Xi
摘要
Autism spectrum disorder (ASD) encompasses a range of neurodevelopmental conditions marked by challenges in social interaction, communication, and behavior. Currently, there are no specific medications for ASD; instead, early identification and intervention are crucial for enhancing brain function and reducing the disorder's negative impacts through specialized education and rehabilitation. Identification of ASD can be achieved through two primary methods. The first is a manual approach, which involves identifying the disorder by observing the individual or interviewing a parent or caregiver. This approach is laborintensive and subjective. The second method employs automated techniques using both traditional machine learning (ML) and advanced deep learning (DL) strategies, which analyze facial images. The human face, reflecting brain conditions, serves as a practical biomarker for early identification. However, identification of ASD from facial images presents significant challenges due to inter-class similarities, intra-class differences, and variations in facial poses. No existing research comprehensively addresses all these issues within a unified model. In this paper, we introduce a novel multi-feature, multi-level cross-fusion transformer network (MFMLCFT), designed to effectively address these challenges. The proposed method enables effective fusion between facial landmark features and multiple image features. Additionally, it focuses on global and local attention features, learning the multi-level relationships among local regions, and combines global and local features to efficiently discern ASD-related facial features. Specifically, our model first applies spatial attention to multiple image features to accentuate salient facial regions, followed by a multi-level cross-fusion transformer encoder that fuses landmark and image features, capturing the internal dynamics of local facial regions and the interactions between global and local features. The outputs from multiple transformer encoders yield a detailed profile of ASD-related facial features. Furthermore, a vision transformer network is integrated to further enhance feature assimilation. Comprehensive experimental results on an ASD facial image dataset demonstrate that our model significantly outperforms state-of-the-art methods, achieving high metrics of accuracy (96.79%), precision (97.12%), recall (96.43%), F1 score (96.77%), and AUC (99.49). This model offers a valuable tool for healthcare professionals to confirm initial screenings and identify children with ASD more effectively.