印地语
误传
计算机科学
人工智能
自然语言处理
粗俗
心理学
显著性(神经科学)
情绪分析
水准点(测量)
社会化媒体
可读性
造谣
作者
Raghvendra Kumar,Pulkit Bansal,Raunak Kumar Singh,Sriparna Saha
标识
DOI:10.1109/taffc.2025.3649246
摘要
Misinformation poses a growing threat across media ecosystems, yet research on Hindi, one of the world's most widely spoken languages, remains limited. We introduce a novel multimodal Hindi dataset of 6,544 article–image pairs to advance misinformation detection. Unlike existing English-centric and predominantly unimodal datasets, ours integrates text, images, and affective signals while being carefully cleaned of veracity cues to avoid artefact-driven inflation. Each sample is annotated with sentiment and emotions, making this the first Hindi resource with multimodal and affective dimensions. Through extensive experiments using IndicBART, IndicBERT, mBERT, and Vision Transformer models, we demonstrate the effectiveness of text–image fusion and affective features across multiple configurations. We also analyze the readability characteristics of genuine and misleading articles, providing insights into the linguistic patterns of Hindi misinformation. This dataset establishes a robust benchmark for multimodal misinformation detection and lays essential groundwork for research in Hindi and other low-resource languages.
科研通智能强力驱动
Strongly Powered by AbleSci AI