计算机科学
突出
模式识别(心理学)
面部表情
人工智能
可视化
变压器
面部表情识别
编码器
面部识别系统
量子力学
操作系统
物理
电压
作者
Xiaohan Xia,Dongmei Jiang
标识
DOI:10.1016/j.ins.2023.119301
摘要
Facial expression recognition rarely explores complex spatiotemporal dependencies among facial regions at different scales. This paper proposes a transformer-based three-layer hierarchical architecture that incorporates multi-scale spatiotemporal aggregation for dynamic facial expression recognition. The hierarchical structure consists of bottom-to-top layers, each comprising transformer encoders with local self-attention mechanisms. These encoders gradually expand their receptive fields through hierarchical spatiotemporal aggregation, enabling the modeling of spatiotemporal context dependencies among facial regions at different scales and across consecutive frames. Consequently, the bottom-to-top layers correspond to learning the fine-grained, coarse-grained, and global facial representations. To evaluate the performance of our proposed framework, we conducted extensive experiments on four public datasets. The comparison results demonstrate that our proposed framework outperforms the state-of-the-art, with accuracies of 79.09%, 62.19%, 64.85%, and 59.79% on the RML, eNTERFACE'05, RAVDESS, and AFEW datasets, respectively. Ablation experiments, statistical significance tests, and visualization analyses indicate that the proposed framework successfully learns emotional-salient facial representations.
科研通智能强力驱动
Strongly Powered by AbleSci AI