Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research

Reduced, Reused and Recycled: The Life of a Dataset in Machine Learning Research
复制标题

DOI:
--
复制
发表时间:
2021-12
期刊:
ArXiv
影响因子:
--
通讯作者:
Bernard Koch;Emily L. Denton;A. Hanna;J. Foster
Bernard Koch;Emily L. Denton;A. Hanna;J. Foster
中科院分区:
其他
文献类型:
--
作者:
Bernard Koch;Emily L. Denton;A. Hanna;J. Foster

文献摘要

被引文献

相似文献

基准数据集在机器学习研究的组织中发挥着核心作用。它们协调研究人员围绕共同的研究问题,并作为实现共同目标的进展的衡量标准。尽管基准测试实践在这一领域发挥着基础性作用,但在机器学习子社区内部或跨机器学习子社区,对基准数据集使用和重用的动态关注相对较少。在本文中,我们将深入研究这些动态。我们研究了2015-2020年机器学习子社区和时间段内数据集使用模式的差异。我们发现越来越多的集中在任务社区内越来越少的数据集,从其他任务的数据集的显着采用,并集中在整个领域的数据集,已介绍了位于少数精英机构的研究人员。我们的研究结果对科学评估,人工智能伦理和该领域内的公平/准入具有影响。
Benchmark datasets play a central role in the organization of machine learning research. They coordinate researchers around shared research problems and serve as a measure of progress towards shared goals. Despite the foundational role of benchmarking practices in this field, relatively little attention has been paid to the dynamics of benchmark dataset use and reuse, within or across machine learning subcommunities. In this paper, we dig into these dynamics. We study how dataset usage patterns differ across machine learning subcommunities and across time from 2015-2020. We find increasing concentration on fewer and fewer datasets within task communities, significant adoption of datasets from other tasks, and concentration across the field on datasets that have been introduced by researchers situated within a small number of elite institutions. Our results have implications for scientific evaluation, AI ethics, and equity/access within the field.