Hate‐UDF: Explainable Hateful Meme Detection With Uncertainty‐Aware Dynamic Fusion

计算机科学 融合 哲学 语言学
作者
Xia Lei,Siqi Wang,Yongkai Fan,Wenqian Shang
出处
期刊:Software - Practice and Experience [Wiley]
卷期号:55 (5): 883-895
标识
DOI:10.1002/spe.3403
摘要

ABSTRACT Background With the increasing integration of Artificial Intelligence (AI) and Internet of Things (IoT), the dissemination of multimodal data is undergoing revolutionary changes. To mitigate the societal risks posed by the rapid spread of malicious multimodal data, such as hateful memes, it is crucial to develop effective detection methods for such data. Existing detection models often struggle with data quality issues and lack interpretability, limiting their effectiveness in content moderation tasks. Aims This paper aims to propose an explainable hateful meme detection model by uncertainty‐aware dynamic fusion. The goal is to enhance both generalization performance and interpretability, addressing the limitations of conventional static fusion methods and existing algorithms for hateful meme detection. Materials & Methods To mitigate the societal risks posed by the rapid spread of malicious multimodal data, such as hateful memes, it is crucial to develop effective detection methods for such data. However, existing algorithms for hateful meme detection frequently overlook the data quality and the interpretability of model. To adress these challenges, this paper proposes Hate‐UDF, an explainable hateful meme detection model with uncertainty‐aware dynamic fusion, providing both high generalization ability and interpretability. This method dynamically evaluates the uncertainty of different modalities, obtains dynamic weights, and utilizes them to weight the feature values for fusion, thereby obtaining a uncertainty‐aware dynamic fusion method with provable upper bounds on generalization error. Furthermore, an analysis of the dynamic weights can explain the modality on which the model primarily relies for detection, thereby providing a method that is both explainable and reliable. Results We compare the performance of Hate‐UDF with three general models and three State of the Art (SOTA) models in the field of hateful meme detection on the Facebook Hateful Memes (FHM) and the Multimedia Automatic Misogyny Identification (MAMI) datasets. Hate‐UDF achieved state‐of‐the‐art performance, surpassing existing models on both datasets. Specifically, it improved accuracy and AUC by 7.56% and 2.8% on FHM and by 3.34% and 0.17% on MAMI compared with the current SOTA model, respectively. Additionally, we demonstrate that the visual modality is more important than the textual modality in the hateful meme detection model, and we explain the primary reason behind this by visualization. Discussion The model dynamically adapts to modality quality, enhancing reliability and reducing the risk of misclassification. Its interpretability, achieved through visualizations of modality and feature attributions, provides valuable insights for content moderation systems and highlights the importance of image modality in detecting hateful meme. While Hate‐UDF provides an explainable and reliable method for detecting hateful memes, it may still learn biases from the training data, potentially leading to the over‐detection of content from certain groups or communities. Future research must focus on improving the fairness and ethical responsibilities of the model's decisions. Conclusion This paper introduces the model of Hate‐UDF, a dynamic fusion method based on uncertainty, designed to improve multimodal fusion issues in existing hateful meme detection models. The model determines the reliability of different modal information by assessing their uncertainty and generates dynamic weights accordingly. By comparing these weights, the model can identify which modality is most influential in detecting malicious content. Therefore, the Hate‐UDF model not only has interpretability but also its generalization performance has been validated.
最长约 10秒,即可获得该文献文件

科研通智能强力驱动
Strongly Powered by AbleSci AI
科研通是完全免费的文献互助平台,具备全网最快的应助速度,最高的求助完成率。 对每一个文献求助,科研通都将尽心尽力,给求助人一个满意的交代。
实时播报
刚刚
dilong完成签到,获得积分10
刚刚
萌羊发布了新的文献求助10
2秒前
2秒前
丘比特应助快毕业采纳,获得10
3秒前
3秒前
彭于晏应助ZhouKunlu采纳,获得10
3秒前
3秒前
3秒前
4秒前
xiaopang完成签到,获得积分10
5秒前
sciexplorer完成签到,获得积分10
5秒前
5秒前
kunnao完成签到,获得积分10
5秒前
ngyx完成签到 ,获得积分10
6秒前
传奇3应助dick_zhang采纳,获得30
6秒前
6秒前
糊涂的萍完成签到,获得积分10
7秒前
doraemon发布了新的文献求助10
7秒前
8秒前
Liu完成签到,获得积分20
8秒前
8秒前
8秒前
baozibaozi完成签到,获得积分10
8秒前
raindrop完成签到,获得积分10
8秒前
活力的番茄完成签到,获得积分10
8秒前
不以完成签到,获得积分10
9秒前
wzyshzu发布了新的文献求助10
9秒前
9秒前
9秒前
没烦恼发布了新的文献求助10
9秒前
GDMUcyc发布了新的文献求助10
10秒前
10秒前
糊涂的萍发布了新的文献求助20
10秒前
haierke完成签到 ,获得积分10
11秒前
aprise发布了新的文献求助10
11秒前
科研通AI2S应助YiYi采纳,获得10
12秒前
wang发布了新的文献求助10
12秒前
12秒前
13秒前
高分求助中
(应助此贴封号)【重要!!请各用户(尤其是新用户)详细阅读】【科研通的精品贴汇总】 10000
Principles of town planning: translating concepts to applications 1000
Cognitive Psychology in a Changing World 800
内視鏡的に摘除しえた十二指腸乳頭部腫瘍の2例 660
Management and the Arts 510
Matrix Methods in Data Mining and Pattern Recognition Second Edition 510
The Neuroscience of Language 400
热门求助领域 (近24小时)
化学 材料科学 医学 生物 纳米技术 工程类 有机化学 化学工程 生物化学 计算机科学 内科学 物理 复合材料 催化作用 细胞生物学 无机化学 光电子学 物理化学 电极 基因
热门帖子
关注 科研通微信公众号,转发送积分 7683701
求助须知:如何正确求助?哪些是违规求助? 9247456
关于积分的说明 19947023
捐赠科研通 7256492
什么是DOI,文献DOI怎么找? 3288495
关于科研通互助平台的介绍 2445841
邀请新用户注册赠送积分活动 2292541