Fair active learning

Fair active learning
复制标题

公平主动学习

DOI:
10.1016/j.eswa.2022.116981
复制
发表时间:
2022
影响因子:
8.5
通讯作者:
Thirumuruganathan, Saravanan
Thirumuruganathan, Saravanan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Anahideh, Hadis;Asudeh, Abolfazl;Thirumuruganathan, Saravanan

文献摘要

相似文献

机器学习(ML)正越来越多地被用于影响社会的高风险应用程序。因此,ML模型不传播歧视是至关重要的。在社交应用中收集准确的标签数据具有挑战性且成本高昂。主动学习是一种很有前途的方法,通过在标签预算内交互查询甲骨文来构建准确的分类器。我们引入公平的主动学习框架来仔细选择要标记的数据点,以平衡模型的准确性和公平性。为了在主动学习采样核中引入公平的概念,需要在添加每个未标记样本后测量模型的公平性。由于它们的标签事先是未知的,我们提出了一个期望公平度量来概率地衡量每个样本对每个可能的类标签的影响。接下来,我们提出了多种优化方案来平衡精度和公平性之间的平衡。我们的第一个优化使用一个控制参数将期望的公平性与熵进行线性聚合。为了避免对期望公平性的错误估计,我们提出了一种嵌套的方法来保持模型的精度,将搜索空间限制在具有大熵的样本点的顶部桶。最后,为了保证标注后模型的不公平性降低,我们提出了复制那些真正降低标注后不公平性的点。我们在广泛使用的基准数据集上,使用人口平等性和公平的均衡赔率概念,证明了我们所提出的算法的有效性和效率。
Machine learning (ML) is increasingly being used in high-stakes applications impacting society. Therefore, it is of critical importance that ML models do not propagate discrimination. Collecting accurate labeled data in societal applications is challenging and costly. Active learning is a promising approach to build an accurate classifier by interactively querying an oracle within a labeling budget. We introduce the fair active learning framework to carefully select data points to be labeled so as to balance model accuracy and fairness. To incorporate the notion of fairness in the active learning sampling core, it is required to measure the fairness of the model after adding each unlabeled sample. Since their labels are unknown in advance, we propose an expected fairness metric to probabilistically measure the impact of each sample if added for each possible class label. Next, we propose multiple optimizations to balance the trade-off between accuracy and fairness. Our first optimization linearly aggregate the expected fairness with entropy using a control parameter. To avoid erroneous estimation of the expected fairness, we propose a nested approach to maintain the accuracy of the model, limiting the search space to the top bucket of sample points with large entropy. Finally, to ensure the unfairness reduction of the model after labeling, we propose to replicate the points that truly reduce the unfairness after labeling. We demonstrate the effectiveness and efficiency of our proposed algorithms over widely used benchmark datasets using demographic parity and equalized odds notions of fairness.