Dataset Geography: Mapping Language Data to Language Users

Dataset Geography: Mapping Language Data to Language Users
复制标题

DOI:
10.18653/v1/2022.acl-long.239
复制
发表时间:
2021-12
期刊:
--
影响因子:
--
通讯作者:
FAHIM FAISAL;Yinkai Wang;Antonios Anastasopoulos
FAHIM FAISAL;Yinkai Wang;Antonios Anastasopoulos
中科院分区:
其他
文献类型:
--
作者:
FAHIM FAISAL;Yinkai Wang;Antonios Anastasopoulos

文献摘要

被引文献

相似文献

随着语言技术变得越来越普遍,人们越来越努力地扩大自然语言处理(NLP)系统的语言多样性和覆盖范围。可以说,影响现代NLP系统质量的最重要因素是数据可用性。在这项工作中,我们研究了NLP数据集的地理代表性,旨在量化NLP数据集是否以及在多大程度上符合语言使用者的预期需求。在此过程中,我们使用实体识别和链接系统,还对它们的跨语言一致性进行了重要的观察,并为更强大的评估提供了建议。最后,我们探讨了一些地理和经济因素,可以解释观察到的数据集分布。
As language technologies become more ubiquitous, there are increasing efforts towards expanding the language diversity and coverage of natural language processing (NLP) systems. Arguably, the most important factor influencing the quality of modern NLP systems is data availability. In this work, we study the geographical representativeness of NLP datasets, aiming to quantify if and by how much do NLP datasets match the expected needs of the language speakers. In doing so, we use entity recognition and linking systems, also making important observations about their cross-lingual consistency and giving suggestions for more robust evaluation. Last, we explore some geographical and economic factors that may explain the observed dataset distributions.