话语
计算机科学
水准点(测量)
人工智能
人工神经网络
图形
代表(政治)
光学(聚焦)
机器学习
深度学习
模式识别(心理学)
自然语言处理
理论计算机科学
地图学
物理
光学
政治
政治学
法学
地理
作者
Shiwen Zhao,Yanjiao Zhang,Yikai Su,Kaifeng Su,Jiemin Liu,Tao Wang,Shiqi Yu
出处
期刊:Sensors
[Multidisciplinary Digital Publishing Institute]
日期:2025-07-21
卷期号:25 (14): 4520-4520
摘要
The global prevalence of depression necessitates the application of technological solutions, particularly sensor-based systems, to augment scarce resources for early diagnostic purposes. In this study, we use benchmark datasets that contain multimodal data including video, audio, and transcribed text. To address depression detection as a chronic long-term disorder reflected by temporal behavioral patterns, we propose a novel framework that segments videos into utterance-level instances using GRU for contextual representation, and then constructs graphs where utterance embeddings serve as nodes connected through dual relationships capturing both chronological development and intermittent relevant information. Graph neural networks are employed to learn multi-dimensional edge relationships and align multimodal representations across different temporal dependencies. Our approach achieves superior performance with an MAE of 5.25 and RMSE of 6.75 on AVEC2014, and CCC of 0.554 and RMSE of 4.61 on AVEC2019, demonstrating significant improvements over existing methods that focus primarily on momentary expressions.
科研通智能强力驱动
Strongly Powered by AbleSci AI