Prompting for Multimodal Hateful Meme Classification

Prompting for Multimodal Hateful Meme Classification
复制标题

提示进行多模式仇恨模因分类

DOI:
--
复制
发表时间:
2023
期刊:
Conference on Empirical Methods in Natural Language Processing
影响因子:
--
通讯作者:
Jing Jiang
Jing Jiang
中科院分区:
--
文献类型:
--
作者:
Rui Cao;R. Lee;Wen;Jing Jiang

文献摘要

被引文献

相似文献

仇恨模因分类是一个具有挑战性的多模态任务,需要复杂的推理和上下文背景知识。理想情况下,我们可以利用明确的外部知识库来补充仇恨模因中的上下文和文化信息。然而,没有已知的明确的外部知识库,可以提供这样的仇恨言论上下文信息。为了解决这一差距,我们提出了一个简单而有效的基于仇恨的模型,它提示预训练的语言模型(PLM)进行仇恨模因分类。具体来说,我们构造简单的提示,并提供了一些上下文中的例子,利用隐含的知识在预先训练的RoberTa语言模型仇恨模因分类。我们在两个公开的仇恨和攻击性模因数据集上进行了广泛的实验。我们的实验结果表明,AdvertHate能够达到90.96的高AUC,在仇恨模因分类任务上优于最先进的基线。我们还进行了细粒度的分析和案例研究的各种提示设置,并证明了仇恨模因分类的提示的有效性。
Hateful meme classification is a challenging multimodal task that requires complex reasoning and contextual background knowledge. Ideally, we could leverage an explicit external knowledge base to supplement contextual and cultural information in hateful memes. However, there is no known explicit external knowledge base that could provide such hate speech contextual information. To address this gap, we propose PromptHate, a simple yet effective prompt-based model that prompts pre-trained language models (PLMs) for hateful meme classification. Specifically, we construct simple prompts and provide a few in-context examples to exploit the implicit knowledge in the pre-trained RoBERTa language model for hateful meme classification. We conduct extensive experiments on two publicly available hateful and offensive meme datasets. Our experiment results show that PromptHate is able to achieve a high AUC of 90.96, outperforming state-of-the-art baselines on the hateful meme classification task. We also perform fine-grain analyses and case studies on various prompt settings and demonstrate the effectiveness of the prompts on hateful meme classification.