Mitigating dataset harms requires stewardship: Lessons from 1000 papers

Mitigating dataset harms requires stewardship: Lessons from 1000 papers
复制标题

DOI:
--
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Kenny Peng;Arunesh Mathur;Arvind Narayanan
Kenny Peng;Arunesh Mathur;Arvind Narayanan
中科院分区:
其他
文献类型:
--
作者:
Kenny Peng;Arunesh Mathur;Arvind Narayanan

文献摘要

相似文献

机器学习数据集引发了人们对隐私、偏见和不道德应用的担忧,导致DukeMTMC、MS-Celeb-1M和Tiny Images等著名数据集被撤回。对此,机器学习界呼吁提高数据集创建的道德标准。为了帮助为这些努力提供信息,我们研究了三个有影响力但存在伦理问题的人脸和个人识别数据集--标签为Face in the Wild(LFW)、MS-Celeb-1M和DukeMTM--通过分析近1000篇引用它们的论文。我们发现,衍生数据集和模型的创建、更广泛的技术和社会变化、许可证缺乏清晰度以及数据集管理实践可能会引发广泛的伦理问题。最后,我们建议了一种分布式的危害缓解方法,该方法考虑了数据集的整个生命周期。
Machine learning datasets have elicited concerns about privacy, bias, and unethical applications, leading to the retraction of prominent datasets such as DukeMTMC, MS-Celeb-1M, and Tiny Images. In response, the machine learning community has called for higher ethical standards in dataset creation. To help inform these efforts, we studied three influential but ethically problematic face and person recognition datasets -- Labeled Faces in the Wild (LFW), MS-Celeb-1M, and DukeMTM -- by analyzing nearly 1000 papers that cite them. We found that the creation of derivative datasets and models, broader technological and social change, the lack of clarity of licenses, and dataset management practices can introduce a wide range of ethical concerns. We conclude by suggesting a distributed approach to harm mitigation that considers the entire life cycle of a dataset.