情态动词
计算机科学
人工智能
多模态
萧条(经济学)
模式识别(心理学)
化学
万维网
宏观经济学
经济
高分子化学
作者
Yaowei Wang,Zulong Lin,Yan Teng,Yuqi Cheng,Haixia Jiang,Yun Yang
标识
DOI:10.1109/tcss.2025.3542986
摘要
Research indicates significant differences in the voice and facial expressions of individuals with depression compared to healthy individuals. Consequently, many studies have begun using audio-visual data for automatic depression detection (ADD). Despite significant progress, numerous challenges remain. During feature extraction, many studies fail to capture crucial spatiotemporal information for depression diagnosis, leading to decreased feature quality. In the modality fusion stage, many methods assume that the semantic information of different modalities is temporally aligned, which is not the case. To address the first challenge, this study proposes a linear spatiotemporal detector (LSTD) to model spatiotemporal information through detection and ensemble, enhancing feature quality. To tackle the second challenge, a cross-modal temporal aligner (CMTA) is proposed, using cross-modal attention for alignment before modality fusion to promote thorough fusion. LSTD and CMTA form the proposed spatiotemporal information modeling and modal alignment (SIMMA) framework. Extensive experiments on five depression datasets (AVEC2013, AVEC2014, AVEC2017, AVEC2019, and CMDep) demonstrate that our method outperforms previous studies in effectiveness.
科研通智能强力驱动
Strongly Powered by AbleSci AI