Word embeddings are biased. But whose bias are they reflecting?

Word embeddings are biased. But whose bias are they reflecting?
复制标题

词嵌入是有偏差的。

DOI:
10.1007/s00146-022-01443-w
复制
发表时间:
2022
期刊:
影响因子:
3
通讯作者:
Ibrahim C. Hashim
Ibrahim C. Hashim
中科院分区:
--
文献类型:
--
作者:
Davor Petreski;Ibrahim C. Hashim

文献摘要

参考文献

被引文献

相似文献

从简历解析到网络搜索和推荐系统,word2vec和其他单词嵌入技术越来越多地出现在人类社会的日常互动中。偏见,如性别偏见,已经得到了深入的研究,并被证明存在于单词嵌入中。大多数研究集中于在向量空间本身的框架内发现和减轻性别偏见。然而,谁的偏向反映在词的嵌入上还没有得到调查。除了发现和减轻性别偏见,同样重要的是要检查词嵌入的偏见中是否体现了以女性为中心的观点或男性中心的观点。这样,我们不仅可以更深入地了解上述偏误的来源,而且还可以为研究自然语言处理系统中的偏误提供一种新的方法。基于之前的社会科学研究和性别研究,我们假设,以男性为中心,或被称为男性中心的偏见在单词嵌入中占主导地位。为了验证这一假设,我们使用了公开可用的最大的英语单词联想测试数据集。我们在词嵌入向量空间中比较了男性和女性参与者对线索词的反应距离。我们发现,单词嵌入偏向于以男性为中心的观点,主要反映了单词联想测试数据集中男性参与者的世界观。因此,通过进行这项研究,我们的目标是揭示在检查算法中的公平性时需要考虑的另一层偏见。
From Curriculum Vitae parsing to web search and recommendation systems, Word2Vec and other word embedding techniques have an increasing presence in everyday interactions in human society. Biases, such as gender bias, have been thoroughly researched and evidenced to be present in word embeddings. Most of the research focuses on discovering and mitigating gender bias within the frames of the vector space itself. Nevertheless, whose bias is reflected in word embeddings has not yet been investigated. Besides discovering and mitigating gender bias, it is also important to examine whether a feminine or a masculine-centric view is represented in the biases of word embeddings. This way, we will not only gain more insight into the origins of the before mentioned biases, but also present a novel approach to investigating biases in Natural Language Processing systems. Based on previous research in the social sciences and gender studies, we hypothesize that masculine-centric, otherwise known as androcentric, biases are dominant in word embeddings. To test this hypothesis we used the largest English word association test data set publicly available. We compare the distance of the responses of male and female participants to cue words in a word embedding vector space. We found that the word embedding is biased towards a masculine-centric viewpoint, predominantly reflecting the worldviews of the male participants in the word association test data set. Therefore, by conducting this research, we aimed to unravel another layer of bias to be considered when examining fairness in algorithms.
DOI: 10.1073/pnas.1720347115
发表时间: 2018-04-17
影响因子: 11.1
作者:
Garg, Nikhil;Schiebinger, Londa;Zou, James
通讯作者: Zou, James
黑人之于罪犯就像白人之于警察:检测和消除词嵌入中的多类偏差
DOI: --
发表时间: 2019
期刊: 2019 Annual Conference of the North American Chapter of the Association for Computational Linguistics (NAACL
影响因子: --
作者:
Manzini, Thomas;Lim, Yao Chong;Tsvetkov, Yulia;Black, Alan W
通讯作者: Black, Alan W
DOI: 10.18653/v1/n18-2003
发表时间: 2018-04
期刊: ArXiv
影响因子: --
作者:
Jieyu Zhao;Tianlu Wang;Mark Yatskar;Vicente Ordonez;Kai-Wei Chang
通讯作者: Jieyu Zhao;Tianlu Wang;Mark Yatskar;Vicente Ordonez;Kai-Wei Chang