Fairness-aware Model-agnostic Positive and Unlabeled Learning

Fairness-aware Model-agnostic Positive and Unlabeled Learning
复制标题

DOI:
10.1145/3531146.3533225
复制
发表时间:
2022-06
期刊:
Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency
影响因子:
--
通讯作者:
Ziwei Wu;Jingrui He
Ziwei Wu;Jingrui He
中科院分区:
其他
文献类型:
--
作者:
Ziwei Wu;Jingrui He

文献摘要

相似文献

随着机器学习在高风险决策问题中的应用日益增多,对某些社会群体的潜在算法偏见对个人和整个社会都产生了负面影响。在现实场景中,许多此类问题涉及正例和无标记数据,比如医学诊断、犯罪风险评估和推荐系统。例如,在医学诊断中,只有确诊的疾病会被记录(正例),其他的则不会(无标记)。尽管在(半)监督和无监督环境下,关于具有公平意识的机器学习已有大量工作,但在上述正例无标记学习(PUL)背景下,公平性问题在很大程度上尚未得到充分研究,而在此背景下该问题通常更为严重。在本文中,为了缓解这种紧张局面,我们提出了一种具有公平意识的PUL方法,名为FairPUL。特别是,对于来自两个群体的个体的二分类问题,我们旨在使两个群体的真阳性率和假阳性率相似,以此作为我们的公平性度量标准。基于对PUL的最优公平分类器的分析,我们设计了一个与模型无关的后处理框架,同时利用正例和无标记例。我们的框架在分类误差和公平性度量方面都被证明在统计上是一致的。在合成数据集和真实数据集上的实验表明,我们的框架在PUL和公平分类方面都优于现有技术水平。
With the increasing application of machine learning in high-stake decision-making problems, potential algorithmic bias towards people from certain social groups poses negative impacts on individuals and our society at large. In the real-world scenario, many such problems involve positive and unlabeled data such as medical diagnosis, criminal risk assessment and recommender systems. For instance, in medical diagnosis, only the diagnosed diseases will be recorded (positive) while others will not (unlabeled). Despite the large amount of existing work on fairness-aware machine learning in the (semi-)supervised and unsupervised settings, the fairness issue is largely under-explored in the aforementioned Positive and Unlabeled Learning (PUL) context, where it is usually more severe. In this paper, to alleviate this tension, we propose a fairness-aware PUL method named FairPUL. In particular, for binary classification over individuals from two populations, we aim to achieve similar true positive rates and false positive rates in both populations as our fairness metric. Based on the analysis of the optimal fair classifier for PUL, we design a model-agnostic post-processing framework, leveraging both the positive examples and unlabeled ones. Our framework is proven to be statistically consistent in terms of both the classification error and the fairness metric. Experiments on the synthetic and real-world data sets demonstrate that our framework outperforms state-of-the-art in both PUL and fair classification.