Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech

Generate, Prune, Select: A Pipeline for Counterspeech Generation against Online Hate Speech
复制标题

DOI:
10.18653/v1/2021.findings-acl.12
复制
发表时间:
2021-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Wanzheng Zhu;S. Bhat
Wanzheng Zhu;S. Bhat
中科院分区:
其他
文献类型:
--
作者:
Wanzheng Zhu;S. Bhat

文献摘要

相似文献

在不阻碍言论自由的情况下有效打击网上日益增长的仇恨言论的对策具有重大的社会利益。自然语言生成(NLG)是开发可扩展解决方案的唯一能力。然而,现成的NLG方法主要是顺序对顺序的神经模型,它们的局限性在于它们产生常见的、重复的和安全的反应,而不考虑仇恨言论(例如,“请不要使用这种语言”)。或无关紧要的回应,使它们对缓和令人憎恨的对话无效。在本文中,我们设计了一种三模块流水线的方法,有效地提高了多样性和相关性。我们提出的流水线首先通过生成模型生成各种反语候选以促进多样性,然后使用BERT模型过滤不符合语法的候选,最后使用一种新的基于检索的方法来选择最相关的反语响应。在三个具有代表性的数据集上的大量实验证明了该方法在生成多样化且相关的反语方面的有效性。
Countermeasures to effectively fight the ever increasing hate speech online without blocking freedom of speech is of great social interest. Natural Language Generation (NLG), is uniquely capable of developing scalable solutions. However, off-the-shelf NLG methods are primarily sequence-to-sequence neural models and they are limited in that they generate commonplace, repetitive and safe responses regardless of the hate speech (e.g.,"Please refrain from using such language.") or irrelevant responses, making them ineffective for de-escalating hateful conversations. In this paper, we design a three-module pipeline approach to effectively improve the diversity and relevance. Our proposed pipeline first generates various counterspeech candidates by a generative model to promote diversity, then filters the ungrammatical ones using a BERT model, and finally selects the most relevant counterspeech response using a novel retrieval-based method. Extensive Experiments on three representative datasets demonstrate the efficacy of our approach in generating diverse and relevant counterspeech.