The Dataset Multiplicity Problem: How Unreliable Data Impacts Predictions

The Dataset Multiplicity Problem: How Unreliable Data Impacts Predictions
复制标题

数据集多重性问题:不可靠的数据如何影响预测

DOI:
10.1145/3593013.3593988
复制
发表时间:
2023
期刊:
ACM
影响因子:
--
通讯作者:
D'Antoni, Loris
D'Antoni, Loris
中科院分区:
--
文献类型:
--
作者:
Meyer, Anna P.;Albarghouthi, Aws;D'Antoni, Loris

文献摘要

参考文献

被引文献

相似文献

我们引入了数据集多样性,这是一种研究训练数据集中的不准确性、不确定性和社会偏见如何影响测试时间预测的方法。数据集多重性框架提出了一个反事实的问题,即如果我们能够以某种方式访问数据集的所有假设的、无偏的版本,那么结果模型集(以及相关的测试时间预测)将是什么。我们讨论了如何使用这个框架来封装数据集真实性中的各种不确定性来源,包括系统性社会偏见,数据收集实践和嘈杂的标签或特征。我们展示了如何准确分析数据集多样性对特定模型架构和不确定性类型的影响:具有标签错误的线性模型。我们的实证分析表明,在合理的假设下,真实世界的数据集包含许多测试样本,其预测受到数据集多重性的影响。此外,特定领域数据集多重性定义的选择决定了哪些样本受到影响,以及不同的人口统计群体是否受到影响。最后,我们讨论了数据集多样性对机器学习实践和研究的影响,包括模型结果不应被信任的考虑因素。
We introduce dataset multiplicity, a way to study how inaccuracies, uncertainty, and social bias in training datasets impact test-time predictions. The dataset multiplicity framework asks a counterfactual question of what the set of resultant models (and associated test-time predictions) would be if we could somehow access all hypothetical, unbiased versions of the dataset. We discuss how to use this framework to encapsulate various sources of uncertainty in datasets’ factualness, including systemic social bias, data collection practices, and noisy labels or features. We show how to exactly analyze the impacts of dataset multiplicity for a specific model architecture and type of uncertainty: linear models with label errors. Our empirical analysis shows that real-world datasets, under reasonable assumptions, contain many test samples whose predictions are affected by dataset multiplicity. Furthermore, the choice of domain-specific dataset multiplicity definition determines what samples are affected, and whether different demographic groups are disparately impacted. Finally, we discuss implications of dataset multiplicity for machine learning practice and research, including considerations for when model outcomes should not be trusted.
DOI: --
发表时间: 2020
期刊: International Conference on Machine Learning
影响因子: --
作者:
Sanghamitra Dutta;Dennis Wei;Hazar Yueksel;Pin;Sijia Liu;Kush R. Varshney
通讯作者: Kush R. Varshney
在选择性标签下描述一组好模型的公平性
DOI: --
发表时间: 2021
期刊: International Conference on Machine Learning
影响因子: --
作者:
Amanda Coston;Ashesh Rambachan;A. Chouldechova
通讯作者: A. Chouldechova
算法单一文化与社会福利
影响因子: 11.1
作者:
J. Kleinberg;Manish Raghavan
通讯作者: Manish Raghavan
DOI: --
发表时间: 2021-08
期刊: --
影响因子: --
作者:
Frances Ding;Moritz Hardt;John Miller;Ludwig Schmidt
通讯作者: Frances Ding;Moritz Hardt;John Miller;Ludwig Schmidt
DOI: 10.1609/aaai.v32i1.11610
发表时间: 2018-01
期刊: ArXiv
影响因子: --
作者:
Xuezhou Zhang;Xiaojin Zhu;Stephen J. Wright
通讯作者: Xuezhou Zhang;Xiaojin Zhu;Stephen J. Wright