计算机科学
多模态
接头(建筑物)
变压器
情态动词
模式治疗法
语音识别
多模式学习
多通道交互
人工智能
自然语言处理
工程类
人机交互
电气工程
电压
结构工程
心理学
万维网
化学
高分子化学
心理治疗师
作者
Kuijie Zhang,Hongjuan Pei,Chenkai Zhang,Shanchen Pang
标识
DOI:10.1109/tce.2025.3593333
摘要
Understanding user sentiment from multiple modalities is essential for applications such as personalized recommendations and affective computing. We propose TMJR, a Three-modal Joint Representation model for multimodal sentiment analysis, which introduces a structured three-round cross-modal attention mechanism to integrate textual, acoustic, and visual information progressively. Unlike prior one-pass or pairwise fusion methods, TMJR models multiple directional interaction paths to capture deeper cross-modal dependencies and improve semantic alignment. We evaluate TMJR on two benchmark datasets: CMU-MOSI and CMU-MOSEI. It achieves state-of-the-art results, with an F1 score of 86.4% and a Pearson correlation of 0.799 on MOSI, and an F1 score of 86.5% and a correlation of 0.775 on MOSEI. Further ablation, label-efficiency, and visualization analyses confirm the model’s robustness. TMJR provides a promising solution for sentiment-aware recommendation and multimodal user understanding.
科研通智能强力驱动
Strongly Powered by AbleSci AI