Representation Bias in Data: A Survey on Identification and Resolution Techniques

Representation Bias in Data: A Survey on Identification and Resolution Techniques
复制标题

DOI:
10.1145/3588433
复制
发表时间:
2022-03
影响因子:
16.6
通讯作者:
N. Shahbazi;Yin Lin;Abolfazl Asudeh;H. V. Jagadish
N. Shahbazi;Yin Lin;Abolfazl Asudeh;H. V. Jagadish
中科院分区:
计算机科学1区
文献类型:
--
作者:
N. Shahbazi;Yin Lin;Abolfazl Asudeh;H. V. Jagadish

文献摘要

相似文献

数据驱动的算法仅与他们使用的数据一样好,而数据集(尤其是社交数据)通常无法适当地代表数据的偏见,这是由于各种原因而发生的。数据采集​​和准备方法。在机器学习模型中的广泛研究(包括几个审查论文)的偏见虽然偏见较少,但本文的研究却较少。以后消耗的范围。多个设计尺寸,并对它们的属性进行并排比较。在他们各自的领域。
Data-driven algorithms are only as good as the data they work with, while datasets, especially social data, often fail to represent minorities adequately. Representation Bias in data can happen due to various reasons, ranging from historical discrimination to selection and sampling biases in the data acquisition and preparation methods. Given that “bias in, bias out,” one cannot expect AI-based solutions to have equitable outcomes for societal applications, without addressing issues such as representation bias. While there has been extensive study of fairness in machine learning models, including several review papers, bias in the data has been less studied. This article reviews the literature on identifying and resolving representation bias as a feature of a dataset, independent of how consumed later. The scope of this survey is bounded to structured (tabular) and unstructured (e.g., image, text, graph) data. It presents taxonomies to categorize the studied techniques based on multiple design dimensions and provides a side-by-side comparison of their properties. There is still a long way to fully address representation bias issues in data. The authors hope that this survey motivates researchers to approach these challenges in the future by observing existing work within their respective domains.