Word embeddings quantify 100 years of gender and ethnic stereotypes

Word embeddings quantify 100 years of gender and ethnic stereotypes
复制标题

DOI:
10.1073/pnas.1720347115
复制
发表时间:
2018-04-17
影响因子:
11.1
通讯作者:
Zou, James
Zou, James
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Garg, Nikhil;Schiebinger, Londa;Zou, James

文献摘要

被引文献

相似文献

词嵌入是一个强大的机器学习框架,它通过一个向量来表示每个英语单词。这些向量之间的几何关系捕获了相应单词之间有意义的语义关系。在本文中,我们开发了一个框架,以展示如何嵌入的时间动态有助于量化的变化,在20世纪和21世纪的美国对妇女和少数民族的刻板印象和态度。我们将在100 y文本数据上训练的词嵌入与美国人口普查相结合,以表明嵌入的变化与人口统计和职业变化密切相关。嵌入捕捉社会的变化,例如,20世纪60年代的妇女运动和亚洲移民到美国,也阐明了具体的形容词和职业如何随着时间的推移与某些人群联系得更紧密。我们对词嵌入的时间分析框架在机器学习和定量社会科学之间开辟了一个富有成效的交叉点。
Word embeddings are a powerful machine-learning framework that represents each English word by a vector. The geometric relationship between these vectors captures meaningful semantic relationships between the corresponding words. In this paper, we develop a framework to demonstrate how the temporal dynamics of the embedding helps to quantify changes in stereotypes and attitudes toward women and ethnic minorities in the 20th and 21st centuries in the United States. We integrate word embeddings trained on 100 y of text data with the US Census to show that changes in the embedding track closely with demographic and occupation shifts over time. The embedding captures societal shifts-e.g., the women's movement in the 1960s and Asian immigration into the United States-and also illuminates how specific adjectives and occupations became more closely associated with certain populations over time. Our framework for temporal analysis of word embedding opens up a fruitful intersection between machine learning and quantitative social science.