When Fair Classification Meets Noisy Protected Attributes

When Fair Classification Meets Noisy Protected Attributes
复制标题

DOI:
10.1145/3600211.3604707
复制
发表时间:
2023-07
期刊:
Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society
影响因子:
--
通讯作者:
Avijit Ghosh;Pablo Kvitca;Chris L. Wilson
Avijit Ghosh;Pablo Kvitca;Chris L. Wilson
中科院分区:
其他
文献类型:
--
作者:
Avijit Ghosh;Pablo Kvitca;Chris L. Wilson

文献摘要

被引文献

相似文献

算法公平性的可操作性带来了一些实际挑战,其中最重要的是数据集中受保护属性的可用性或可靠性。在现实世界中,实际和法律的障碍可能会阻止人口统计数据的收集和使用,从而难以确保算法的公平性。虽然最初的公平算法没有考虑这些限制,最近的建议旨在通过将噪声在受保护的属性或不使用受保护的属性在所有实现算法的公平性分类。据我们所知,这是公平分类算法的第一个头对头的研究,比较属性依赖,噪声容忍和属性不知道的算法沿着预测性和公平性的双轴。我们通过对四个真实世界数据集和合成扰动的案例研究来评估这些算法。我们的研究表明,属性不知道和噪声容忍公平分类器可以实现类似的性能水平的属性依赖算法,即使受保护的属性是嘈杂的。然而,在实践中实施这些建议需要仔细的细微差别。我们的研究提供了深入了解使用公平的分类算法的实际影响的情况下,受保护的属性是嘈杂的或部分可用的。
The operationalization of algorithmic fairness comes with several practical challenges, not the least of which is the availability or reliability of protected attributes in datasets. In real-world contexts, practical and legal impediments may prevent the collection and use of demographic data, making it difficult to ensure algorithmic fairness. While initial fairness algorithms did not consider these limitations, recent proposals aim to achieve algorithmic fairness in classification by incorporating noisiness in protected attributes or not using protected attributes at all. To the best of our knowledge, this is the first head-to-head study of fair classification algorithms to compare attribute-reliant, noise-tolerant and attribute-unaware algorithms along the dual axes of predictivity and fairness. We evaluated these algorithms via case studies on four real-world datasets and synthetic perturbations. Our study reveals that attribute-unaware and noise-tolerant fair classifiers can potentially achieve similar level of performance as attribute-reliant algorithms, even when protected attributes are noisy. However, implementing them in practice requires careful nuance. Our study provides insights into the practical implications of using fair classification algorithms in scenarios where protected attributes are noisy or partially available.