计算机科学
水准点(测量)
稳健性(进化)
人工智能
机器学习
特征提取
钥匙(锁)
特征(语言学)
情绪识别
噪音(视频)
航程(航空)
点(几何)
情感计算
情绪分类
活动识别
多模态
人机交互
任务分析
评价方法
作者
Zheng Lian,Licai Sun,Yong Ren,Hao Gu,Haiyang Sun,Lan Chen,Bin Liu,Jianhua Tao
标识
DOI:10.1109/tpami.2026.3653457
摘要
Multimodal emotion recognition plays a vital role in enhancing user experience in human-computer interaction. Over the past few decades, researchers have developed a range of algorithms and made remarkable progress. While each approach demonstrates certain advantages, inconsistent choices in feature extraction methods, evaluation protocols, and experimental settings have hindered fair comparisons among them. These inconsistencies significantly impede the advancement of the field. To address this issue, we introduce MERBench, a unified evaluation benchmark for multimodal emotion recognition. Our goal is to assess the contributions of several key techniques commonly used in prior studies, such as feature selection, multimodal fusion, robustness analysis, fine-tuning, and pre-training. We believe this work offers clear and comprehensive guidance for future research. Based on the evaluation results of MERBench, we further point out some promising research directions. In addition, we present a new emotion dataset, MER2023, specifically designed for the Chinese language environment. This dataset serves as a benchmark for research in multi-label learning, noise robustness, and semi-supervised learning.
科研通智能强力驱动
Strongly Powered by AbleSci AI