自动汇总
计算机科学
社会化媒体
编码器
自然语言处理
变压器
人工智能
资源(消歧)
情报检索
数据科学
机器学习
万维网
量子力学
电压
操作系统
计算机网络
物理
作者
Jawaher Alghamdi,Yuqing Lin,Suhuai Luo
标识
DOI:10.1016/j.knosys.2024.111884
摘要
The proliferation of fake news across languages and domains on social media platforms poses a significant societal threat. Current automatic detection methods for low-resource languages (e.g., Swahili, Indonesian and other low-resource languages) face limitations due to two factors: sequential length restrictions in pre-trained language models (PLMs) like multilingual bidirectional encoder representation from transformers (mBERT), and the presence of noisy training data. This work proposes a novel and efficient multilingual fake news detection (MFND) approach that addresses these challenges. Our solution leverages a hybrid extractive and abstractive summarization strategy to extract only the most relevant content from news articles. This significantly reduces data length while preserving crucial information for fake news classification. The pre-processed data is then fed into mBERT for classification. Extensive evaluations on a publicly available multilingual dataset demonstrate the superiority of our approach compared to state-of-the-art (SOTA) methods. Our analysis, both quantitative and qualitative, highlights the strengths of this method, achieving new performance benchmarks and emphasizing the impact of content condensation on model accuracy and efficiency. This framework paves the way for faster, more accurate MFND, fostering more robust information ecosystems.
科研通智能强力驱动
Strongly Powered by AbleSci AI