Discovering and Categorising Language Biases in Reddit

Discovering and Categorising Language Biases in Reddit
复制标题

DOI:
10.1609/icwsm.v15i1.18048
复制
发表时间:
2020-08
期刊:
--
影响因子:
--
通讯作者:
Xavier Ferrer Aran;T. Nuenen;J. Such;N. Criado
Xavier Ferrer Aran;T. Nuenen;J. Such;N. Criado
中科院分区:
其他
文献类型:
--
作者:
Xavier Ferrer Aran;T. Nuenen;J. Such;N. Criado

文献摘要

相似文献

我们提出了一种数据驱动的方法,使用词嵌入来发现和分类讨论平台Reddit上的语言偏见。作为孤立的用户社区的空间,Reddit等平台越来越多地与种族主义、性别歧视和其他形式的歧视问题联系在一起,这表明有必要监测这些群体的语言。在大型文本数据集中跟踪语言偏见的最有前途的人工智能方法之一涉及词嵌入,它将文本转换为高维密集向量并捕获词之间的语义关系。然而,以前的研究需要预先定义的潜在偏倚集来研究,例如,性别是否或多或少与特定类型的工作有关。这使得这些方法不适合处理较小的以社区为中心的数据集,例如Reddit上的数据集,其中包含较小的词汇和俚语,以及可能特定于该社区的偏见。本文提出了一种数据驱动的方法来自动发现Reddit上在线话语社区词汇中编码的语言偏见。在我们的方法中,受保护的属性连接到数据中发现的评估词,然后通过语义分析系统进行分类。我们通过比较我们在Google News数据集中发现的偏差与以前文献中发现的偏差来验证我们方法的有效性。然后,我们成功地发现了不同Reddit社区中的性别偏见,宗教偏见和种族偏见。最后,我们讨论了这种数据驱动的偏差发现方法的潜在应用场景和局限性。
We present a data-driven approach using word embeddings to discover and categorise language biases on the discussion platform Reddit. As spaces for isolated user communities, platforms such as Reddit are increasingly connected to issues of racism, sexism and other forms of discrimination, signalling the need to monitor the language of these groups. One of the most promising AI approaches to trace linguistic biases in large textual datasets involves word embeddings, which transform text into high-dimensional dense vectors and capture semantic relations between words. Yet, previous studies require predefined sets of potential biases to study, e.g., whether gender is more or less associated with particular types of jobs. This makes these approaches unfit to deal with smaller and community-centric datasets such as those on Reddit, which contain smaller vocabularies and slang, as well as biases that may be particular to that community. This paper proposes a data-driven approach to automatically discover language biases encoded in the vocabulary of online discourse communities on Reddit. In our approach, protected attributes are connected to evaluative words found in the data, which are then categorised through a semantic analysis system. We verify the effectiveness of our method by comparing the biases we discover in the Google News dataset with those found in previous literature. We then successfully discover gender bias, religion bias, and ethnic bias in different Reddit communities. We conclude by discussing potential application scenarios and limitations of this data-driven bias discovery method.