可解释性
杠杆(统计)
计算机科学
人工智能
情绪分析
模式
机器学习
保险丝(电气)
编码
深度学习
标记数据
深层神经网络
自然语言处理
数据挖掘
作者
Chenguang Song,Chao Ke,Bingjing Jia,Yiqing Shen
标识
DOI:10.1038/s41598-025-19850-6
摘要
Recent advances in sentiment analysis have primarily focused on fusing multimodal information from video data, including visual, acoustic, and textual features, across temporal sequences. While great effort has been made to integrate or fuse information across modalities, less is known about the extent to which temporal segments contribute to model decisions. In addition, current interpretable methods, such as prototype networks, are primarily designed for uni-modal analysis and fail to handle the complex interactions between multiple modalities and temporal dependencies inherent in video data. To address the challenges, we propose MultiModal Prototypical Networks (MMPNet), which extends prototype-based interpretability to multimodal sentiment classification. Specifically, MMPNet can identify contributions of time-level features and leverage them to explain why a particular prediction was made, while also helping to find the relative importance of modality-level features. Experimental results show that MMPNet outperforms existing methods by 2.9% and 1.6% in accuracy on CMU-MOSI and CMU-MOSEI respectively, and achieves better interpretability.
科研通智能强力驱动
Strongly Powered by AbleSci AI