概化理论
计算机科学
人工智能
稳健性(进化)
可视化
遮罩(插图)
模式识别(心理学)
代表(政治)
计算机视觉
特征提取
情感计算
人工神经网络
图像(数学)
像素
情绪检测
情绪识别
相关性(法律)
机器学习
对比度(视觉)
数据建模
认知心理学
深度学习
情绪分类
视觉掩蔽
分层数据库模型
语音识别
上下文模型
注意力网络
数据可视化
心理学
作者
Weiye Peng,Sheng-hua Zhong,Ahmed Fares,Yan Liu
标识
DOI:10.1109/taffc.2025.3641966
摘要
Compared to conventional image content analysis tasks, visual emotion analysis is perceived as a complex, abstract, and potentially culturally dependent endeavor. The accuracy of automatic image-based emotion recognition remains a challenge, and the most significant obstacles are the affective gap and scarcity of data, particularly labeled data. To address these challenge, this paper proposes a saliency-guided masked image modeling approach. Specifically, the proposed framework employs multi-modal large model to generate more emotional images, thereby reducing the impact of data scarcity on model performance. Subsequently, neuroimaging and behavioral studies have demonstrated that human visual attention is attracted by the emotional relevance of a stimulus. In light of this, our model employs a saliency-guided masking strategy to identify emotion-related regions for masking sampling to fit the affective gap. In contrast to the conventional approach of using the original pixel values for the reconstruction target, our model eliminates high-frequency components from the pixels, thus enhancing the generalizability of the model. The use of this unsupervised representation learning approach enables the model to exhibit outstanding recognition performance in downstream emotion recognition tasks on three standard emotion datasets. Furthermore, ablation experiments, robustness test, and visualization experiments corroborate the effectiveness of the proposed method.
科研通智能强力驱动
Strongly Powered by AbleSci AI