可解释性
计算机科学
人工智能
预处理器
软件部署
特征(语言学)
推论
模式识别(心理学)
计算机视觉
特征提取
面子(社会学概念)
面部识别系统
身份(音乐)
深度学习
面部表情
机器学习
人脸检测
芯(光纤)
目标检测
特征向量
干扰(通信)
作者
Liang Zhang,Peng Zhang,Tianhuan Huang,Jian Zhao,Wei Xiang,Xianye Ben
标识
DOI:10.1109/jiot.2026.3667276
摘要
Automatic depression detection from facial videos is promising for IoT deployment but has long suffered from the black-box nature of deep models, which overlook how depression manifests on the face. This lack of interpretability leads to redundant feature learning, overly complex architectures, and consequently a trade-off between accuracy and deployability. We introduce a lightweight, interpretable system that explicitly models static–dynamic facial patterns, color channels, and regional textures. At its core is a Hybrid Static–Dynamic Model (HSDM) with detachable modules and identity decoupling, supported by two task-driven preprocessing steps (blue-channel suppression and 90° Gabor filtering). On AVEC2013/2014, our approach reduces mean absolute error by 10.1% on AVEC2014 while maintaining competitive results on AVEC2013. From a systems perspective, it achieves real-time inference on Jetson Nano (0.071M parameters, 0.21 GFLOPs, 17.69 ms latency, 56.54 samples/s throughput), outperforming larger SOTA models under identical conditions. The design facilitates on-device deployment and offers transparent decision paths through interpretable static–dynamic fusion, color synergy, and region-level analysis.
科研通智能强力驱动
Strongly Powered by AbleSci AI