作者
Lang He,Weizhao Yang,Junnan Zhao,Haifeng Chen,Dongmei Jiang
摘要
Major depressive disorder (MDD) is projected to become one of the leading mental disorders by 2030. While audiovisual cues have garnered significant attention in depression recognition research owing to their non-invasive acquisition and rich emotional expressiveness. However, conventional centralized training paradigms raise substantial privacy concerns for individuals with depression and are further hindered by data heterogeneity and label inconsistency across datasets. To overcome these challenges, a hybrid architecture, termed Federated Domain Adversarial with Attention Mechanism (FedDAAM), for privacy preserving multimodal depression assessment, is proposed. FedDAAM introduces a mechanism by differentiating discriminative features into depression-public and depression-private features. Specifically, to extract visual depression-private features from the AVEC2013 and AVEC2014 datasets, a local attention-aware (LAA) architecture is developed. For the depression-public features, action units (AUs), landmarks, head poses, and eye gazes features are adopted. In addition, to consider the transferability and performance of individual client, a dynamic parameter aggregation mechanism, termed FedDyA, is proposed. Extensive validations are performed on the AVEC2013, AVEC2014 and AVEC2017 databases, resulting in root mean square error (RMSE) and mean absolute error (MAE) of 8.61/6.78, 8.59/6.77, and 4.71/3.68, respectively. More importantly, to the best of our knowledge, this is the first study to borrow federated learning (FL) for multimodal depression assessment. The proposed framework offers a novel solution for privacy-aware, distributed clinical diagnosis of depression. Code will be available at: https://github.com/helang818/FedDAAM/.