Counterfactual Fairness in Synthetic Data Generation

Counterfactual Fairness in Synthetic Data Generation
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Mahed Abroshan;Mohammad Mahdi Khalili;†. AndrewElliott
Mahed Abroshan;Mohammad Mahdi Khalili;†. AndrewElliott
中科院分区:
其他
文献类型:
--
作者:
Mahed Abroshan;Mohammad Mahdi Khalili;†. AndrewElliott

文献摘要

被引文献

相似文献

由于隐私问题,综合数据生成(SDG)被提议作为数据共享的有前途的数据共享解决方案,因此释放真实数据集并非一个选择。虽然私人可持续发展目标的主要目标是创建一个数据集,以保留对数据集有贡献的个人的隐私,但合成数据的使用也为改善源头上的公平性问题提供了机会。由于数据集中存在历史偏见,因此使用偏置数据训练ML模型可能会导致不公平的模型,这可能会加剧歧视。使用合成数据,我们可以尝试在释放数据之前从数据集中删除偏差。在这项工作中,我们将合成数据生成中公平性的定义形式化,并提出一种实现反事实公平的方法。
Synthetic data generation (SDG) is proposed as a promising solution for data sharing as in many high-stake applications due to privacy concerns, releasing the real dataset is not an option. While the main goal of private SDG is to create a dataset that preserves the privacy of individuals contributing to the dataset, the use of synthetic data also creates an opportunity to improve the fairness issue at the source. Since there exist historical biases in the datasets, using the biased data to train an ML model can lead to an unfair model which may exacerbate the discrimination. Using synthetic data, we can attempt to remove the bias from the dataset before releasing the data. In this work, we formalize the definition of fairness in synthetic data generation and propose a method to achieve counterfactual fairness.