AAEBERT: Debiasing BERT-based Hate Speech Detection Models via Adversarial Learning

AAEBERT: Debiasing BERT-based Hate Speech Detection Models via Adversarial Learning
复制标题

AAEBERT:通过对抗性学习消除基于 BERT 的仇恨言论检测模型的偏差

DOI:
--
复制
发表时间:
2022
期刊:
International Conference on Machine Learning and Applications
影响因子:
--
通讯作者:
Feng Luo
Feng Luo
中科院分区:
--
文献类型:
--
作者:
Ebuka Okpala;Long Cheng;N. Mbwambo;Feng Luo

文献摘要

参考文献

相似文献

仇恨语音数据集包含机器学习模型传播的偏见。当这些模型对用非裔美国人英语(AAE)写的推文进行分类时,他们预测AAE推文被认为是仇恨/辱骂的比率高于用标准美国英语(SAE)写的推文。本文评估了为仇恨语音检测而微调的语言模型中的偏差,以及对抗性学习在减少此类偏差方面的有效性。我们介绍了AAEBERT,这是一个针对非裔美国人英语的预训练语言模型,它是在AAE推文的基础上对Bert进行重新训练而得到的。AAEBERT用于提取各种仇恨语音数据集中每条推文的表示,并将推文分为两类-AAE方言和非AAE方言。以AAEBERT的表示和方言标签为输入的三层前馈神经网络作为去偏的对抗性网络。我们评估了为仇恨言论检测而微调的语言模型中的偏差。然后通过比较应用对抗性去偏向前后的结果来评估这些模型中对抗性去偏向的有效性。分析表明,微调后的模型偏向于AAE,对抗性去偏能有效地降低偏倚。
Hate speech datasets contain bias which machine learning models propagate. When these models classify tweets written in African American English (AAE), they predict AAE tweets as hate/abusive at a higher rate than tweets written in Standard American English (SAE). This paper assesses bias in language models fine-tuned for hate speech detection and the effectiveness of adversarial learning in reducing such bias. We introduce AAEBERT, a pre-trained language model for African American English obtained by re-training BERT-base on AAE tweets. AAEBERT is used to extract the representation of each tweet in the various hate speech datasets and to classify tweets into two classes - AAE dialect and non-AAE dialect. A three-layer feedforward neural network that takes the representation from AAEBERT and a dialect label as input is used as the adversarial network for debiasing. We evaluate bias in language models fine-tuned for hate speech detection. Then assess the effectiveness of adversarial debiasing in these models by comparing results before and after adversarial debiasing is applied. Analysis reveals that the fine-tuned models are biased towards AAE, and adversarial debiasing is effective in reducing bias.
DOI: 10.1002/poi3.85
发表时间: 2015-06-01
影响因子: 4.9
作者:
Burnap, Pete;Williams, Matthew L.
通讯作者: Williams, Matthew L.