BBQ: A hand-built bias benchmark for question answering

BBQ: A hand-built bias benchmark for question answering
复制标题

DOI:
10.18653/v1/2022.findings-acl.165
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
Alicia Parrish;Angelica Chen;Nikita Nangia;Vishakh Padmakumar;Jason Phang;Jana Thompson;Phu Mon Htut;Sam Bowman
Alicia Parrish;Angelica Chen;Nikita Nangia;Vishakh Padmakumar;Jason Phang;Jana Thompson;Phu Mon Htut;Sam Bowman
中科院分区:
其他
文献类型:
--
作者:
Alicia Parrish;Angelica Chen;Nikita Nangia;Vishakh Padmakumar;Jason Phang;Jana Thompson;Phu Mon Htut;Sam Bowman

文献摘要

被引文献

相似文献

NLP模型学习社会偏见是有据可查的,但关于这些偏见如何在问答(QA)等应用任务的模型输出中表现出来的工作却很少。我们介绍了偏见基准QA(烧烤),一个数据集的问题集,由作者构建,突出证明社会偏见对人属于受保护的类沿着九个社会层面相关的美国英语环境。我们的任务在两个层面上评估模型的反应:(i)在信息不足的情况下,我们测试反应如何强烈地反映社会偏见,以及(ii)在信息充足的情况下,我们测试模型的偏见是否覆盖正确的答案选择。我们发现,当上下文信息不足时,模型通常依赖于刻板印象,这意味着模型的输出在这种情况下始终再现有害的偏见。虽然当上下文提供了一个信息丰富的答案时,模型更准确,但它们仍然依赖于刻板印象,当正确答案与社会偏见相一致时的准确性平均高出3.4个百分点,而不是当它冲突时,这种差异扩大到针对性别的例子超过5个百分点。
It is well documented that NLP models learn social biases, but little work has been done on how these biases manifest in model outputs for applied tasks like question answering (QA). We introduce the Bias Benchmark for QA (BBQ), a dataset of question-sets constructed by the authors that highlight attested social biases against people belonging to protected classes along nine social dimensions relevant for U.S. English-speaking contexts. Our task evaluate model responses at two levels: (i) given an under-informative context, we test how strongly responses reflect social biases, and (ii) given an adequately informative context, we test whether the model’s biases override a correct answer choice. We find that models often rely on stereotypes when the context is under-informative, meaning the model’s outputs consistently reproduce harmful biases in this setting. Though models are more accurate when the context provides an informative answer, they still rely on stereotypes and average up to 3.4 percentage points higher accuracy when the correct answer aligns with a social bias than when it conflicts, with this difference widening to over 5 points on examples targeting gender for most models tested.