Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic Data

Robin Hood and Matthew Effects: Differential Privacy Has Disparate Impact on Synthetic Data
复制标题

DOI:
--
复制
发表时间:
2021-09
期刊:
--
影响因子:
--
通讯作者:
Georgi Ganev;Bristena Oprisanu;Emiliano De Cristofaro
Georgi Ganev;Bristena Oprisanu;Emiliano De Cristofaro
中科院分区:
其他
文献类型:
--
作者:
Georgi Ganev;Bristena Oprisanu;Emiliano De Cristofaro

文献摘要

被引文献

相似文献

使用差异隐私(DP)训练的生成模型可用于生成合成数据,同时将隐私风险降至最低。我们分析了DP对这些模型的影响,特别是研究了:1)合成数据中类/子组的大小,2)在它们上运行的分类任务的准确性。我们还评估了不同水平的不平衡和隐私预算的影响。我们的分析使用了三个最先进的DP模型(PrivBayes、DP-WGAN和Pate-GaN),并表明DP在生成的合成数据中产生了相反的尺寸分布。它影响多数和少数阶级/小组之间的差距;在某些情况下,通过缩小差距(“罗宾汉”效应),在另一些情况下,通过扩大差距(“马太”效应)。无论哪种方式,这都会对合成数据的分类任务的准确性产生(类似的)不同影响,对数据中未被充分代表的子部分的影响更大。因此,当在合成数据上训练模型时,可能会招致对不同子总体不同对待的风险,导致不可靠或不公平的结论。
Generative models trained with Differential Privacy (DP) can be used to generate synthetic data while minimizing privacy risks. We analyze the impact of DP on these models vis-a-vis underrepresented classes/subgroups of data, specifically, studying: 1) the size of classes/subgroups in the synthetic data, and 2) the accuracy of classification tasks run on them. We also evaluate the effect of various levels of imbalance and privacy budgets. Our analysis uses three state-of-the-art DP models (PrivBayes, DP-WGAN, and PATE-GAN) and shows that DP yields opposite size distributions in the generated synthetic data. It affects the gap between the majority and minority classes/subgroups; in some cases by reducing it (a"Robin Hood"effect) and, in others, by increasing it (a"Matthew"effect). Either way, this leads to (similar) disparate impacts on the accuracy of classification tasks on the synthetic data, affecting disproportionately more the underrepresented subparts of the data. Consequently, when training models on synthetic data, one might incur the risk of treating different subpopulations unevenly, leading to unreliable or unfair conclusions.