Towards Fair Classifiers Without Sensitive Attributes: Exploring Biases in Related Features

Towards Fair Classifiers Without Sensitive Attributes: Exploring Biases in Related Features
复制标题

DOI:
10.1145/3488560.3498493
复制
发表时间:
2021-04
期刊:
Proceedings of the Fifteenth ACM International Conference on Web Search and Data Mining
影响因子:
--
通讯作者:
Tianxiang Zhao;Enyan Dai;Kai Shu;Suhang Wang
Tianxiang Zhao;Enyan Dai;Kai Shu;Suhang Wang
中科院分区:
其他
文献类型:
--
作者:
Tianxiang Zhao;Enyan Dai;Kai Shu;Suhang Wang

文献摘要

被引文献

相似文献

尽管机器学习模型发展迅速并取得了巨大成功,但大量研究暴露了它们从训练数据中继承潜在歧视和社会偏见的缺点。这种现象阻碍了它们在高风险应用程序中的采用。因此,已经做出了许多努力来开发公平的机器学习模型。它们中的大多数要求在训练过程中敏感属性可用,以学习公平模型。然而,在实际应用中,由于隐私或法律的问题,获取敏感属性往往是不可行的,这对现有的公平保证策略提出了挑战。虽然每个数据样本的敏感属性是未知的,但我们观察到训练数据中通常存在一些与敏感属性高度相关的非敏感特征,这可以用来减轻偏差。因此,在本文中,我们研究了一个新的问题,探索与敏感属性高度相关的特征,以学习公平和准确的分类器。我们从理论上证明,通过最小化这些相关特征与模型预测之间的相关性,我们可以学习一个公平的分类器。基于这一动机,我们提出了一个新的框架,同时使用这些相关的功能进行准确的预测和执行公平。此外,该模型可以动态调整每个相关特征的正则化权重,以平衡其对模型分类和公平性的贡献。在真实数据集上的实验结果证明了该模型在学习公平模型时的有效性,并具有较高的分类精度。
Despite the rapid development and great success of machine learning models, extensive studies have exposed their disadvantage of inheriting latent discrimination and societal bias from the training data. This phenomenon hinders their adoption on high-stake applications. Thus, many efforts have been taken for developing fair machine learning models. Most of them require that sensitive attributes are available during training to learn fair models. However, in many real-world applications, it is usually infeasible to obtain the sensitive attributes due to privacy or legal issues, which challenges existing fair-ensuring strategies. Though the sensitive attribute of each data sample is unknown, we observe that there are usually some non-sensitive features in the training data that are highly correlated with sensitive attributes, which can be used to alleviate the bias. Therefore, in this paper, we study a novel problem of exploring features that are highly correlated with sensitive attributes for learning fair and accurate classifiers. We theoretically show that by minimizing the correlation between these related features and model prediction, we can learn a fair classifier. Based on this motivation, we propose a novel framework which simultaneously uses these related features for accurate prediction and enforces fairness. In addition, the model can dynamically adjust the regularization weight of each related feature to balance its contribution on model classification and fairness. Experimental results on real-world datasets demonstrate the effectiveness of the proposed model for learning fair models with high classification accuracy.