Fairness and Accuracy Under Domain Generalization

Fairness and Accuracy Under Domain Generalization
复制标题

DOI:
10.48550/arxiv.2301.13323
复制
发表时间:
2023-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Thai-Hoang Pham;Xueru Zhang;Ping Zhang
Thai-Hoang Pham;Xueru Zhang;Ping Zhang
中科院分区:
其他
文献类型:
--
作者:
Thai-Hoang Pham;Xueru Zhang;Ping Zhang

文献摘要

相似文献

随着机器学习(ML)算法越来越多地应用于高风险的应用程序中,人们担心它们可能对某些社会群体存在偏见。尽管已经提出了许多方法来使ML模型公平,但它们通常依赖于训练和部署中的数据分布相同的假设。不幸的是,这一点在实践中经常被违反,一个在培训期间公平的模式可能会在部署期间导致意想不到的结果。尽管数据集平移下的稳健最大似然模型的设计问题已经得到了广泛的研究,但现有的工作大多只关注精度的传递。在本文中,我们研究了在测试时的数据可以从从未见过的域中采样的域泛化情况下,公平性和准确性的转移。我们首先给出了不公平性和部署时期望损失的理论界,然后通过不变表示学习得到了公平性和准确性能够完美传递的充分条件。在此指导下,我们设计了一种学习算法,使得利用训练数据学习的公平ML模型在部署环境变化时仍然具有较高的公平性和准确性。在真实数据上的实验验证了该算法的有效性。模型实施可在https://github.com/pth1993/FATDM.上获得
As machine learning (ML) algorithms are increasingly used in high-stakes applications, concerns have arisen that they may be biased against certain social groups. Although many approaches have been proposed to make ML models fair, they typically rely on the assumption that data distributions in training and deployment are identical. Unfortunately, this is commonly violated in practice and a model that is fair during training may lead to an unexpected outcome during its deployment. Although the problem of designing robust ML models under dataset shifts has been widely studied, most existing works focus only on the transfer of accuracy. In this paper, we study the transfer of both fairness and accuracy under domain generalization where the data at test time may be sampled from never-before-seen domains. We first develop theoretical bounds on the unfairness and expected loss at deployment, and then derive sufficient conditions under which fairness and accuracy can be perfectly transferred via invariant representation learning. Guided by this, we design a learning algorithm such that fair ML models learned with training data still have high fairness and accuracy when deployment environments change. Experiments on real-world data validate the proposed algorithm. Model implementation is available at https://github.com/pth1993/FATDM.