计算机科学
自然语言处理
人工智能
人机交互
语言学
哲学
作者
Haimei Qin,Zhiwei Yang,Chaodong Tong,Lei Jiang
标识
DOI:10.1109/hpcc64274.2024.00148
摘要
Detecting harmful memes plays a crucial role in safeguarding social media users and maintaining a healthy network ecosystem. However, the ability of multimodal large language models(MLLMs) in harmful meme detection tasks remains to be fully developed. Fine-tuning MLLMs is prohibitively expensive due to the large number of parameters, making it essential to leverage these powerful models more efficiently rather than simply fine-tuning them. In this work, we propose GENHMD, a generated rationale-enhanced detection framework, which generates rationales by prompting MLLMs. Our key insight is that MLLMs, pre-trained on massive amounts of data containing rich knowledge, can enhance harmful meme detection by eliciting specialized knowledge. Specifically, inspired by the powerful capacity of MLLMs for text generation and reasoning, we first employ MLLMs to generate credible reasons through chain-of-thought reasoning. To further promote collaboration between harmfulness rationales and the multimodal information inherent in memes, we leverage a smaller multimodal model to learn the attention weights between different modal features using the attention mechanism. Comprehensive experiments on two public meme datasets demonstrate that GENHMD achieves outstanding performance in harmful meme detection. Further analysis shows that the generated rationales substantiate the explainability of the detection results.
科研通智能强力驱动
Strongly Powered by AbleSci AI