Detecting Gender Stereotypes: Lexicon vs. Supervised Learning Methods

Detecting Gender Stereotypes: Lexicon vs. Supervised Learning Methods
复制标题

DOI:
10.1145/3313831.3376488
复制
发表时间:
2020-04
期刊:
Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems
影响因子:
--
通讯作者:
Jenna Cryan;Shiliang Tang;Xinyi Zhang;Miriam J. Metzger;Haitao Zheng;Ben Y. Zhao
Jenna Cryan;Shiliang Tang;Xinyi Zhang;Miriam J. Metzger;Haitao Zheng;Ben Y. Zhao
中科院分区:
其他
文献类型:
--
作者:
Jenna Cryan;Shiliang Tang;Xinyi Zhang;Miriam J. Metzger;Haitao Zheng;Ben Y. Zhao

文献摘要

相似文献

语言的偏见会影响我们之间的互动方式和整个社会。从当今各种情况下,经常观察到语言肯定性别刻板印象,从推荐信和Wikipedia条目到小说小说和电影对话。然而,迄今为止,几乎没有关于用自然语言(特别是英语)量化性别刻板印象的方法的共识。常见方法(包括负责检测性别偏见的公司采用的方法)依赖于1974年的BSRI最初研究的词典方法。在本文中,我们通过相对呈相对呈上的性别刻板印象在现代工具的背景下重新检查性别刻板印象检测的作用分析基于词典的方法和端到端,基于ML的方法在最先进的自然语言处理系统中普遍存在。我们使用大型数据集的努力表明,即使与基于词典的更新方法相比,即使经过适度尺寸的Corpora培训,端到端的分类方法也更加健壮和准确。
Biases in language influence how we interact with each other and society at large. Language affirming gender stereotypes is often observed in various contexts today, from recommendation letters and Wikipedia entries to fiction novels and movie dialogue. Yet to date, there is little agreement on the methodology to quantify gender stereotypes in natural language (specifically the English language). Common methodology (including those adopted by companies tasked with detecting gender bias) rely on a lexicon approach largely based on the original BSRI study from 1974. In this paper, we reexamine the role of gender stereotype detection in the context of modern tools, by comparatively analyzing efficacy of lexicon-based approaches and end-to-end, ML-based approaches prevalent in state-of-the-art natural language processing systems. Our efforts using a large dataset show that even compared to an updated lexicon-based approach, end-to-end classification approaches are significantly more robust and accurate, even when trained by moderately sized corpora.