计算机科学
接头(建筑物)
桥(图论)
培训(气象学)
面部表情识别
面部表情
人工智能
面部识别系统
模式识别(心理学)
动作(物理)
语音识别
表达式(计算机科学)
计算机视觉
工程类
内科学
气象学
物理
建筑工程
程序设计语言
医学
量子力学
作者
Shuyi Mao,Xinpeng Li,Fan Zhang,Xiaojiang Peng,Yang Yang
标识
DOI:10.1109/tmm.2025.3535327
摘要
Label biases in facial expression recognition (FER) datasets, caused by annotators' subjectivity, pose challenges in improving the performance of target datasets when auxiliary labeled data are used. Moreover, training with multiple datasets can lead to visible degradations in the target dataset. To address these issues, we propose a novel framework called the AU-aware Vision Transformer (AU-ViT), which leverages unified action unit (AU) information and discards expression annotations of auxiliary data. AU-ViT integrates an elaborately designed AU branch in the middle part of a master ViT to enhance representation learning during training. Through qualitative and quantitative analyses, we demonstrate that AU-ViT effectively captures expression regions and is robust to real-world occlusions. Additionally, we observe that AU-ViT also yields performance improvements on the target dataset, even without auxiliary data, by utilizing pseudo AU labels. Our AU-ViT achieves performances superior to, or comparable to, that of the state-of-the-art methods on FERPlus, RAFDB, AffectNet, LSD and the other three occlusion test datasets.
科研通智能强力驱动
Strongly Powered by AbleSci AI