阿萨姆人
计算机科学
情绪分析
人工智能
自然语言处理
合并(版本控制)
编码器
判别式
多模态
深度学习
模式识别(心理学)
情报检索
语言学
万维网
操作系统
哲学
作者
Ringki Das,Thoudam Doren Singh
出处
期刊:ACM Transactions on Asian and Low-Resource Language Information Processing
日期:2023-02-17
卷期号:22 (6): 1-30
被引量:33
摘要
Before the arrival of the web as a corpus, people detected positive and negative news based on the understanding of the textual content from physical newspaper rather than an automatic identification approach from readily available e-newspapers. Thus, the earlier sentiment analysis approach is based on unimodal data, and less effort is paid to the multimodal data. However, the presence of multimodal information helps us to get a clearer understanding of the sentiment. To the best of our knowledge, less work has been introduced on the image–text multimodal sentiment analysis framework of Assamese, a low-resource Indian language mostly spoken in the northeast part of India. We built an Assamese news articles dataset consisting of news text and associated images and one image caption to conduct an experimental study. Focusing on important words and discriminative regions of the images mostly related to sentiment, two individual unimodal such as textual and visual models are proposed. The visual model is developed using an encoder-decoder–based image caption generation system. An image–text multimodal approach is proposed to explore the internal correlation between textual and visual features for joint sentiment classification. Finally, we propose the multimodal sentiment analysis framework, i.e., Textual Visual Multimodal Fusion, by employing a late fusion scheme to merge the three different modalities for the final sentiment prediction. Experimental results conducted on the Assamese dataset built in-house demonstrate that the contextual integration of multimodal features delivers better performance than unimodal features.
科研通智能强力驱动
Strongly Powered by AbleSci AI