计算机科学
语音活动检测
人工智能
机器学习
自然语言处理
监督学习
标记数据
语音识别
情绪检测
社会化媒体
语言模型
语音处理
语言习得
训练集
任务分析
深度学习
主动学习(机器学习)
言语共同体
计算语言学
作者
Dongjun Wei,Michael Chau,Zixuan Li
标识
DOI:10.25300/misq/2025/18416
摘要
Hate speech is a major problem on social media platforms. Automatic hate speech detection methods relying on machine learning models, which learn from manually labeled datasets, have been proposed in both academia and industry. However, there is increasing evidence that hate speech detection datasets labeled by general annotators (e.g., amateurs or MTurk workers) contain systematic bias, as they cannot effectively consider language use differences among different speakers. When such biased datasets are used to train machine learning models, the resulting models will also be biased. Unlike general annotators, experts can produce much less biased annotations. However, expert annotations cannot be efficiently obtained in large quantities. This paper bridges the gap by adopting a weakly supervised learning method for hate speech detection using a small number of expert annotations. We propose a novel design that uses contrastive learning and prompt-based learning based on large language models, incorporating a group estimator, a pair generator, and knowledge injection. Using real-world Twitter posts written by African American English speakers and other racial groups as an example, extensive experiments were conducted to demonstrate the superior performance of the proposed method. The proposed approach was also evaluated on data in the LGBTQ+ community and achieved consistent results. The study has important academic and practical implications for hate speech detection and large language models.
科研通智能强力驱动
Strongly Powered by AbleSci AI