计算机科学
人工智能
串联(数学)
水准点(测量)
特征(语言学)
机器学习
代表(政治)
编码器
自编码
冗余(工程)
模式
模式识别(心理学)
模态(人机交互)
独立性(概率论)
特征学习
语义鸿沟
数据挖掘
特征提取
融合
判别式
特征向量
传感器融合
相互信息
深度学习
人工神经网络
条件独立性
相似性(几何)
联营
作者
Shihao Li,Zhuhong Shao,Rongyin Qin,Yongzhen Huang,Peipeng Liang,Xiaobai Li,Yinan Jiang,Yanhe Deng,Tie Liu,Xiaohui Tan
标识
DOI:10.1109/taffc.2025.3611238
摘要
In order to achieve early screening and assist clinical decision-making, automatic depression assessment based on multimodal data are highly anticipated. However, the existed methods often suffer from semantic gap and information redundancy due to heterogeneity among modalities. To address this challenge, this paper investigates a novel Feature Disentanglement and Fusion Network (FDFNet) for predicting depression severity from audio-visual cues. Firstly, we design the shared and private encoders to disentangle modality-shared and modalityprivate representations. The former representation that acquires joint information is subjected by similarity constraints between modalities to ensure their distributions as close as possible. The latter that can capture unique features of each modality is restrained by independence constraints for keeping their distributions distinct. The decoder is then developed to reconstruct unimodal representation with constraints to minimize information loss. Finally, an efficient fusion strategy through addition and concatenation is ultilized for aggregating information. Experimental results on four benchmark datasets demonstrate that the proposed FDFNet consistently outperforms several stateof-the-art methods, with the competitive MAE/RMSE values of 6.22/7.58 on AVEC2013, 5.21/6.49 on AVEC2014, 4.25/5.34 on DAIC-WOZ, and 4.41/5.10 on E-DAIC, indicating that multimodal deep learning based on audio-visual is an attractive solution for objectively evaluating the depression severity.
科研通智能强力驱动
Strongly Powered by AbleSci AI