Self-Supervised Fair Representation Learning without Demographics

Self-Supervised Fair Representation Learning without Demographics
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Junyi Chai;Xiaoqian Wang
Junyi Chai;Xiaoqian Wang
中科院分区:
其他
文献类型:
--
作者:
Junyi Chai;Xiaoqian Wang

文献摘要

被引文献

相似文献

公平性已经成为机器学习中的一个重要话题。一般来说,大多数关于公平性的文献都假设敏感信息(如性别或种族)存在于训练集中,并使用这些信息来减轻偏见。然而,由于隐私和监管等实际问题,这些方法的应用受到限制。此外,尽管许多文献研究了监督学习,但在许多现实场景中,我们希望利用大型未标记数据集来提高模型的准确性。我们能否在没有敏感信息和标签的情况下改进公平分类?为了解决这个问题,在本文中,我们提出了一种新的基于重加权的对比学习方法。我们的方法的目标是在不观察敏感属性的情况下学习一个一般公平的表示。我们的方法根据训练样本相对于验证样本的梯度方向,为每次迭代的训练样本分配权重,以使平均top-k验证损失最小化。与过去没有人口统计的公平方法相比,我们的方法建立在完全无监督的训练数据上,只需要一个小的标记验证集。我们提供了严格的理论证明我们的模型的收敛性。实验结果表明,我们提出的方法实现了更好的或可比的性能比国家的最先进的方法在三个数据集的准确性和几个公平性指标。
Fairness has become an important topic in machine learning. Generally, most literature on fairness assumes that the sensitive information, such as gender or race, is present in the training set, and uses this information to mitigate bias. However, due to practical concerns like privacy and regulation, applications of these methods are restricted. Also, although much of the literature studies supervised learning, in many real-world scenarios, we want to utilize the large unlabelled dataset to improve the model’s accuracy. Can we improve fair classification without sensitive information and without labels? To tackle the problem, in this paper, we propose a novel reweighing-based contrastive learning method. The goal of our method is to learn a generally fair representation without observing sensitive attributes. Our method assigns weights to training samples per iteration based on their gradient directions relative to the validation samples such that the average top-k validation loss is minimized. Compared with past fairness methods without demographics, our method is built on fully unsupervised training data and requires only a small labelled validation set. We provide rigorous theoretical proof of the convergence of our model. Experimental results show that our proposed method achieves better or comparable performance than state-of-the-art methods on three datasets in terms of accuracy and several fairness metrics.