Fair Collective Classification in Networked Data

Fair Collective Classification in Networked Data
复制标题

DOI:
10.1109/bigdata55660.2022.10020610
复制
发表时间:
2022-12
期刊:
2022 IEEE International Conference on Big Data (Big Data)
影响因子:
--
通讯作者:
Karuna Bhaila;Yongkai Wu;Xintao Wu
Karuna Bhaila;Yongkai Wu;Xintao Wu
中科院分区:
其他
文献类型:
--
作者:
Karuna Bhaila;Yongkai Wu;Xintao Wu

文献摘要

相似文献

集体分类通过标签传播利用网络结构信息来提高节点分类任务的预测精度。由于这些模型使用来自先前标记的节点的信息,这些节点通常包含历史偏差,因此它们可能会导致相对于历史偏差的预测。节点的敏感属性,如种族和性别。在整个推理过程中,这种偏差甚至可能由于传播而被放大,特别是对于以同质性为特征的网络。尽管过去和正在进行的公平分类的研究,研究,以确保公平的集体分类仍然是未开发的。在本文中,我们提出了一个公平的集体分类框架(表示为FairCC),并制定各种启发式方法,包括节点重新加权,阈值调整和后处理,以实现公平的预测。我们还实现和测试了几个天真的公平集体分类方法。半合成数据集上的实验突出了朴素方法的不足,并证明了所提出的算法在显着减少预测偏差方面的有效性。
Collective classification utilizes network structure information via label propagation to improve prediction accuracy for node classification tasks. Because these models use information from previously labeled nodes which often contain historical bias, they may result in predictions that are biased w.r.t. the sensitive attributes of nodes such as race and gender. Throughout inference, this bias may even be amplified due to propagation especially for networks characterized by homophily. Despite past and ongoing research on fair classification, research to ensure fair collective classification s till remains unexplored. In this paper, we present a fair collective classification framework (denoted as FairCC) and formulate various heuristic methodologies, including node reweighting, threshold adjustment, and postprocessing, to achieve fair prediction. We also implement and test several naive methodologies for fair collective classification. Experiments on semi-synthetic datasets highlight the insufficiency of the naive methodologies and demonstrate the effectiveness of the proposed heuristics in significantly reducing prediction bias.