Self-Supervised Euphemism Detection and Identification for Content Moderation

Self-Supervised Euphemism Detection and Identification for Content Moderation
复制标题

DOI:
10.1109/sp40001.2021.00075
复制
发表时间:
2021-03
期刊:
2021 IEEE Symposium on Security and Privacy (SP)
影响因子:
--
通讯作者:
Wanzheng Zhu;Hongyu Gong;Rohan Bansal;Zachary Weinberg;Nicolas Christin;G. Fanti;S. Bhat
Wanzheng Zhu;Hongyu Gong;Rohan Bansal;Zachary Weinberg;Nicolas Christin;G. Fanti;S. Bhat
中科院分区:
其他
文献类型:
--
作者:
Wanzheng Zhu;Hongyu Gong;Rohan Bansal;Zachary Weinberg;Nicolas Christin;G. Fanti;S. Bhat

文献摘要

相似文献

边缘团体和组织有着使用委婉语的悠久历史,委婉语是一种具有秘密含义的普通词汇,用来掩盖他们正在讨论的内容。如今,委婉语的一种常见用法是逃避社交媒体平台实施的内容审核政策。现有的执行策略的工具自动依赖于关键字搜索“禁止列表”上的单词,但这些都是众所周知的不精确:即使仅限于脏话,它们仍然会导致令人尴尬的误报[1]。当一个常用的普通词获得了委婉的含义时,将其添加到基于关键词的禁用列表中是无望的:考虑“锅”(存储容器或大麻?)或“加热器”(家用电器或火器?)目前的社交媒体公司雇佣员工手动检查帖子,但这是昂贵的,不人道的,也没有更有效。对于人类版主来说,一个词被委婉地使用通常是显而易见的,但他们可能不知道秘密的含义是什么,因此该消息是否违反政策。此外,当一个委婉语被禁止时,使用它的小组只需要发明另一个,让版主落后一步。本文将展示无监督算法,通过分析单词在其委婉语级别的上下文中,既可以检测单词被委婉地使用,又可以识别每个单词的秘密含义。与现有的使用上下文无关词嵌入的技术相比,我们的委婉语检测算法在文本语料库中对未标记委婉语的检测准确率提高了30-400%。据我们所知,我们揭示单词委婉含义的算法是同类算法中的第一个。在内容版主和政策规避者之间的军备竞赛中,我们的算法可能有助于将平衡转移到版主的方向。
Fringe groups and organizations have a long history of using euphemisms—ordinary-sounding words with a secret meaning—to conceal what they are discussing. Nowadays, one common use of euphemisms is to evade content moderation policies enforced by social media platforms. Existing tools for enforcing policy automatically rely on keyword searches for words on a "ban list", but these are notoriously imprecise: even when limited to swearwords, they can still cause embarrassing false positives [1]. When a commonly used ordinary word acquires a euphemistic meaning, adding it to a keyword-based ban list is hopeless: consider "pot" (storage container or marijuana?) or "heater" (household appliance or firearm?) The current generation of social media companies instead hire staff to check posts manually, but this is expensive, inhumane, and not much more effective. It is usually apparent to a human moderator that a word is being used euphemistically, but they may not know what the secret meaning is, and therefore whether the message violates policy. Also, when a euphemism is banned, the group that used it need only invent another one, leaving moderators one step behind.This paper will demonstrate unsupervised algorithms that, by analyzing words in their sentence-level context, can both detect words being used euphemistically, and identify the secret meaning of each word. Compared to the existing state of the art, which uses context-free word embeddings, our algorithm for detecting euphemisms achieves 30–400% higher detection accuracies of unlabeled euphemisms in a text corpus. Our algorithm for revealing euphemistic meanings of words is the first of its kind, as far as we are aware. In the arms race between content moderators and policy evaders, our algorithms may help shift the balance in the direction of the moderators.