The Risk of Racial Bias in Hate Speech Detection

The Risk of Racial Bias in Hate Speech Detection
复制标题

DOI:
10.18653/v1/p19-1163
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Maarten Sap;Dallas Card;Saadia Gabriel;Yejin Choi;Noah A. Smith
Maarten Sap;Dallas Card;Saadia Gabriel;Yejin Choi;Noah A. Smith
中科院分区:
其他
文献类型:
--
作者:
Maarten Sap;Dallas Card;Saadia Gabriel;Yejin Choi;Noah A. Smith

文献摘要

被引文献

相似文献

我们调查了注释者对方言差异的不敏感性会导致自动仇恨言论检测模型的种族偏见,从而可能会扩大对少数人群的伤害,我们首先发现了非裔美国人英语(AAE)的意外相关性。然后,我们表明,在这些语料库中训练的模型获得并传播这些偏见,因此自我识别的非洲裔美国人的推文和推文与其他人相比,AAE的推文和推文的可能性要高两倍。最后,我们提出 *方言 *和 *种族启动 *作为减少注释中种族偏见的方法,表明当注释者明确意识到AAE Tweet的方言时,它们的可能性大大降低了推文为冒犯性。
We investigate how annotators’ insensitivity to differences in dialect can lead to racial bias in automatic hate speech detection models, potentially amplifying harm against minority populations. We first uncover unexpected correlations between surface markers of African American English (AAE) and ratings of toxicity in several widely-used hate speech datasets. Then, we show that models trained on these corpora acquire and propagate these biases, such that AAE tweets and tweets by self-identified African Americans are up to two times more likely to be labelled as offensive compared to others. Finally, we propose *dialect* and *race priming* as ways to reduce the racial bias in annotation, showing that when annotators are made explicitly aware of an AAE tweet’s dialect they are significantly less likely to label the tweet as offensive.