计算机科学
分类器(UML)
假新闻
可信赖性
社会化媒体
二元分类
注释
人工智能
训练集
透视图(图形)
情报检索
监督学习
万维网
互联网隐私
支持向量机
人工神经网络
作者
Stefan Helmstetter,Heiko Paulheim
标识
DOI:10.1109/asonam.2018.8508520
摘要
The problem of automatic detection of fake news in social media, e.g., on Twitter, has recently drawn some attention. Although, from a technical perspective, it can be regarded as a straight-forward, binary classification problem, the major challenge is the collection of large enough training corpora, since manual annotation of tweets as fake or non-fake news is an expensive and tedious endeavor. In this paper, we discuss a weakly supervised approach, which automatically collects a large-scale, but very noisy training dataset comprising hundreds of thousands of tweets. During collection, we automatically label tweets by their source, i.e., trustworthy or untrustworthy source, and train a classifier on this dataset. We then use that classifier for a different classification target, i.e., the classification of fake and non-fake tweets. Although the labels are not accurate according to the new classification target (not all tweets by an untrustworthy source need to be fake news, and vice versa), we show that despite this unclean inaccurate dataset, it is possible to detect fake news with an F1 score of up to 0.9.
科研通智能强力驱动
Strongly Powered by AbleSci AI