Estimating the success of re-identifications in incomplete datasets using generative models

Estimating the success of re-identifications in incomplete datasets using generative models
复制标题

DOI:
10.1038/s41467-019-10933-3
复制
发表时间:
2019-07-23
影响因子:
16.6
通讯作者:
De Montjoye, Yves-Alexandre
De Montjoye, Yves-Alexandre
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Rocher, Luc;Hendrickx, Julien M.;De Montjoye, Yves-Alexandre

文献摘要

被引文献

相似文献

虽然丰富的医疗、行为和社会人口数据是现代数据驱动研究的关键,但它们的收集和使用引发了合理的隐私担忧。在共享数据集之前,通过去身份识别和采样来匿名化数据集一直是解决这些关切的主要工具。我们在这里提出了一种基于生成性Copula的方法,它可以准确地估计特定人被正确重新识别的可能性,即使在严重不完整的数据集中也是如此。在210个种群上,我们的方法预测个体唯一性的AUC得分在0.84到0.97之间,错误发现率很低。使用我们的模型,我们发现99.98%的美国人可以在使用15个人口统计属性的任何数据集中正确地重新识别。我们的结果表明,即使是高采样的匿名数据集也不太可能满足GDPR提出的现代匿名化标准,并严重挑战去身份发布-遗忘模型的技术和法律充分性。
While rich medical, behavioral, and socio-demographic data are key to modern data-driven research, their collection and use raise legitimate privacy concerns. Anonymizing datasets through de-identification and sampling before sharing them has been the main tool used to address those concerns. We here propose a generative copula-based method that can accurately estimate the likelihood of a specific person to be correctly re-identified, even in a heavily incomplete dataset. On 210 populations, our method obtains AUC scores for predicting individual uniqueness ranging from 0.84 to 0.97, with low false-discovery rate. Using our model, we find that 99.98% of Americans would be correctly re-identified in any dataset using 15 demographic attributes. Our results suggest that even heavily sampled anonymized datasets are unlikely to satisfy the modern standards for anonymization set forth by GDPR and seriously challenge the technical and legal adequacy of the de-identification release-and-forget model.