Auto-Debias: Debiasing Masked Language Models with Automated Biased Prompts

Auto-Debias: Debiasing Masked Language Models with Automated Biased Prompts
复制标题

DOI:
10.18653/v1/2022.acl-long.72
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Yue Guo-;Yi Yang;A. Abbasi
Yue Guo-;Yi Yang;A. Abbasi
中科院分区:
其他
文献类型:
--
作者:
Yue Guo-;Yi Yang;A. Abbasi

文献摘要

被引文献

相似文献

类似人类的偏见和不受欢迎的社会刻板印象存在于大型预训练语言模型中。鉴于这些模型在实际应用中的广泛采用,减轻这种偏差已成为一项新兴的重要任务。在本文中,我们提出了一种自动方法来减轻预训练语言模型中的偏见。与之前使用外部语料库来微调预训练模型的去偏置工作不同,我们通过提示直接探测预训练模型中编码的偏置。具体来说,我们提出了一个变种的光束搜索方法,自动搜索有偏见的提示,这样的完形填空风格的完成是最不同的,相对于不同的人口统计群体。鉴于已识别的有偏见的提示,我们提出了一个分布对齐损失,以减轻偏见。在标准数据集和度量标准上的实验结果表明,我们提出的Auto-Debias方法可以显着减少BERT,RoBERTa和ALBERT等预训练语言模型中的偏见,包括性别和种族偏见。此外,公平性的改善并没有降低语言模型的理解能力,如使用GLUE基准测试所示。
Human-like biases and undesired social stereotypes exist in large pretrained language models. Given the wide adoption of these models in real-world applications, mitigating such biases has become an emerging and important task. In this paper, we propose an automatic method to mitigate the biases in pretrained language models. Different from previous debiasing work that uses external corpora to fine-tune the pretrained models, we instead directly probe the biases encoded in pretrained models through prompts. Specifically, we propose a variant of the beam search method to automatically search for biased prompts such that the cloze-style completions are the most different with respect to different demographic groups. Given the identified biased prompts, we then propose a distribution alignment loss to mitigate the biases. Experiment results on standard datasets and metrics show that our proposed Auto-Debias approach can significantly reduce biases, including gender and racial bias, in pretrained language models such as BERT, RoBERTa and ALBERT. Moreover, the improvement in fairness does not decrease the language models’ understanding abilities, as shown using the GLUE benchmark.