计算机科学
特征(语言学)
人工智能
特征提取
特征学习
背景(考古学)
突出
情绪识别
构造(python库)
相似性(几何)
语义特征
编码器
余弦相似度
钥匙(锁)
情感计算
深度学习
人工神经网络
模态(人机交互)
质量(理念)
机器学习
模式
模式识别(心理学)
情绪分类
编码
多模态
人机交互
原始数据
判别式
面部表情
作者
Jue Feng,Zhengpeng Zhao,Lianmin Zhou,Yuanyuan Pu,Jieping YE,Dan Xu,Jinjing Gu
标识
DOI:10.1109/taffc.2025.3638415
摘要
The quality of features directly affects the accuracy of Multimodal Emotion Recognition (MER). A key challenge in this context is the effective extraction of dynamically interactive multimodal features to enrich conversational emotion representations. However, existing approaches are often constrained by non-end-to-end architectures, overlooking the significance of feature extraction in MER. To address the problem of dynamic interaction in emotion feature extraction, this paper introduces an end-to-end network based on Cross-Modal Hybrid Prompt Learning (CoMPLe). The model takes raw video as input and leverages three prompt mechanisms to guide large-scale pre-trained encoders in extracting emotionally salient features with latent correlations. Specifically, we design a cross-modal soft prompt learning strategy to mine complementary information across modalities and dynamically adjust the cross-modal semantic space. To capture stage-dependent characteristics, deep feature prompts are incorporated to progressively learn intra-modal contextual representations. Furthermore, a label prompt mechanism is proposed to construct hard prompt templates from emotion labels. Finally, the highest cosine similarity is computed between each unimodal feature and the label prompt templates to activate factual knowledge relevant to emotion recognition. Experiments on three public datasets show that the end-to-end network proposed in this paper surpasses the existing State-Of-The-Art baselines.
科研通智能强力驱动
Strongly Powered by AbleSci AI