Feels Bad Man: Dissecting Automated Hateful Meme Detection Through the Lens of Facebook's Challenge

Feels Bad Man: Dissecting Automated Hateful Meme Detection Through the Lens of Facebook's Challenge
复制标题

DOI:
10.36190/2022.65
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Catherine Jennifer;Fatemeh Tahmasbi;Jeremy Blackburn;G. Stringhini;Savvas Zannettou;Emiliano De Cristofaro
Catherine Jennifer;Fatemeh Tahmasbi;Jeremy Blackburn;G. Stringhini;Savvas Zannettou;Emiliano De Cristofaro
中科院分区:
其他
文献类型:
--
作者:
Catherine Jennifer;Fatemeh Tahmasbi;Jeremy Blackburn;G. Stringhini;Savvas Zannettou;Emiliano De Cristofaro

文献摘要

相似文献

网络模因已经成为一种主要的交流方式;然而,与此同时,它们也越来越多地被用来鼓吹极端主义和培养贬损信仰。尽管如此,对于模因的哪些感知方面导致了这种现象,我们并没有一个明确的认识。在这项工作中,我们评估了当前最先进的多模态机器学习模型对仇恨模因检测的有效性,特别是它们在平台上的泛化性。我们使用两个基准数据集,包括来自4chan的“政治不正确”板(/pol/)的12,140和10,567张图片,以及Facebook的仇恨表情包挑战赛数据集,来训练比赛中排名第一的机器学习模型,以发现区分病毒仇恨表情包和良性表情包的最突出特征。我们通过三个实验来确定多模态对分类性能的重要性,边缘Web社区对主流社交平台的影响能力,以及模型在4chan模因上的学习可迁移性。我们的实验表明,模因的图像特征比其文本内容提供了更丰富的信息。我们还发现,目前用于在线检测模因中仇恨言论的系统需要进一步关注其视觉元素,以提高其对潜在文化内涵的解释,这意味着多模态模型未能充分把握模因中仇恨言论的复杂性,也无法在社交媒体平台上进行推广。
Internet memes have become a dominant method of communication; at the same time, however, they are also increasingly being used to advocate extremism and foster derogatory beliefs. Nonetheless, we do not have a firm understanding as to which perceptual aspects of memes cause this phenomenon. In this work, we assess the efficacy of current state-of-the-art multimodal machine learning models toward hateful meme detection, and in particular with respect to their generalizability across platforms. We use two benchmark datasets comprising 12,140 and 10,567 images from 4chan's"Politically Incorrect"board (/pol/) and Facebook's Hateful Memes Challenge dataset to train the competition's top-ranking machine learning models for the discovery of the most prominent features that distinguish viral hateful memes from benign ones. We conduct three experiments to determine the importance of multimodality on classification performance, the influential capacity of fringe Web communities on mainstream social platforms and vice versa, and the models' learning transferability on 4chan memes. Our experiments show that memes' image characteristics provide a greater wealth of information than its textual content. We also find that current systems developed for online detection of hate speech in memes necessitate further concentration on its visual elements to improve their interpretation of underlying cultural connotations, implying that multimodal models fail to adequately grasp the intricacies of hate speech in memes and generalize across social media platforms.