Measuring Social Biases in Grounded Vision and Language Embeddings

Measuring Social Biases in Grounded Vision and Language Embeddings
复制标题

DOI:
10.18653/v1/2021.naacl-main.78
复制
发表时间:
2020-02
期刊:
--
影响因子:
--
通讯作者:
Candace Ross;B. Katz;Andrei Barbu
Candace Ross;B. Katz;Andrei Barbu
中科院分区:
其他
文献类型:
--
作者:
Candace Ross;B. Katz;Andrei Barbu

文献摘要

被引文献

相似文献

我们将测量词嵌入中的社会偏见的概念推广到基于视觉的词嵌入。偏差存在于有根据的嵌入中,并且确实似乎与无根据的嵌入一样或更重要。尽管视觉和语言可能会受到不同的偏见,但人们可能希望这能减轻两者的偏见。有多种方法可以将词嵌入中的度量偏差推广到这个新设置。我们介绍了概括的空间(ground - weat和ground - seat),并证明了三个概括回答了关于偏见、语言和视觉如何相互作用的不同但重要的问题。这些指标被用于一个新的数据集,第一个用于基础偏差,该数据集是通过使用来自COCO、Conceptual Captions和谷歌images的10,228张图像来增强标准语言偏差基准而创建的。数据集的构建是具有挑战性的,因为视觉数据集本身就有很大的偏见。系统中存在的这些偏见将在它们被部署时开始对现实世界产生影响,这使得仔细衡量偏见,然后减轻偏见,对建立一个公平的社会至关重要。
We generalize the notion of measuring social biases in word embeddings to visually grounded word embeddings. Biases are present in grounded embeddings, and indeed seem to be equally or more significant than for ungrounded embeddings. This is despite the fact that vision and language can suffer from different biases, which one might hope could attenuate the biases in both. Multiple ways exist to generalize metrics measuring bias in word embeddings to this new setting. We introduce the space of generalizations (Grounded-WEAT and Grounded-SEAT) and demonstrate that three generalizations answer different yet important questions about how biases, language, and vision interact. These metrics are used on a new dataset, the first for grounded bias, created by augmenting standard linguistic bias benchmarks with 10,228 images from COCO, Conceptual Captions, and Google Images. Dataset construction is challenging because vision datasets are themselves very biased. The presence of these biases in systems will begin to have real-world consequences as they are deployed, making carefully measuring bias and then mitigating it critical to building a fair society.