情绪分析
社会化媒体
计算机科学
人工智能
自然语言处理
深度学习
心理学
数据科学
人机交互
多媒体
微博
情感计算
机器学习
作者
Xianxun Zhu,Heyang Feng,Erik Cambria,Xiaohan Yu,José Santamaría,Xuhui Fan,Rui Wang
标识
DOI:10.1109/taffc.2026.3681216
摘要
Multimodal sentiment analysis (MSA) on social media is increasingly critical for understanding complex emotional expressions, yet it faces significant challenges in data-scarce environments where annotated multimodal content is limited. Here, we present FMSA, a novel framework that integrates prompt-based learning with advanced vision-language models to enable robust few-shot MSA. Our approach leverages an instruction-aware Query Transformer (Q-Former) to dynamically extract and align visual features with task-specific textual prompts, enhancing cross-modal fusion. We introduce a distributed consistency sampling strategy to construct representative few-shot datasets, ensuring statistical diversity under constrained conditions. Evaluated across six benchmark social media datasets-including MVSA-S, MVSA-M, and Twitter-Depression-FMSA outperforms state-of-the-art methods, achieving an accuracy of 63.47% and an F1 score of 57.34% on MVSA-S with full data, and 61.25% accuracy with 56.06% F1 in few-shot settings using just 1% of the data. By fine-tuning lightweight components while preserving pretrained model robustness, FMSA mitigates overfitting and delivers generalizable performance. We also release a curated few-shot dataset as a community resource. This framework advances MSA by offering an efficient, scalable solution for interpreting multimodal emotions in low-resource scenarios.
科研通智能强力驱动
Strongly Powered by AbleSci AI