Application-driven Privacy-preserving Data Publishing with Correlated Attributes

Application-driven Privacy-preserving Data Publishing with Correlated Attributes
复制标题

应用程序驱动的具有相关属性的隐私保护数据发布

DOI:
10.5555/3451271.3451280
复制
发表时间:
2021
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)
影响因子:
--
通讯作者:
Sirajum Munir
Sirajum Munir
中科院分区:
--
文献类型:
--
作者:
A. Rezaei;Chaowei Xiao;Jie Gao;Bo Li;Sirajum Munir

文献摘要

相似文献

计算方面的最新进展使收集有关个人活动和私人生活空间的大量数据成为可能。为了解决用户在这种环境中的隐私问题,我们提出了一种称为PR-GAN的新框架,该框架使用生成对抗网络提供隐私保护机制。给定一个目标应用程序,PR-GAN会自动修改数据以隐藏敏感属性(这些属性可能是隐藏的,可以通过机器学习算法推断出来),同时保留目标应用程序中的数据实用程序。与以前的作品,公众的目标应用程序和敏感属性之间的相关性的可能的知识是内置到我们的建模。我们将我们的问题表述为一个优化问题,证明存在最优解,并使用生成对抗网络(GAN)来创建扰动。我们进一步表明,我们的方法在河豚框架下提供了隐私保证,这是差分隐私的优雅概括,允许对数据和相关性的先验知识进行建模。通过实验,我们表明,我们的方法优于传统的方法,有效地隐藏敏感属性,同时保证高性能的目标应用程序,属性推理和训练的目的。最后,我们通过进一步的实验证明,一旦我们的模型在一组个体上学习了隐私保护任务,例如隐藏受试者的身份,它就可以在一个单独的组上执行相同的任务,并且性能下降最小。
Recent advances in computing have allowed for the possibility to collect large amounts of data on personal activities and private living spaces. To address the privacy concerns of users in this environment, we propose a novel framework called PR-GAN that offers privacy-preserving mechanism using generative adversarial networks. Given a target application, PR-GAN automatically modifies the data to hide sensitive attributes – which may be hidden and can be inferred by machine learning algorithms – while preserving the data utility in the target application. Unlike prior works, the public’s possible knowledge of the correlation between the target application and sensitive attributes is built into our modeling. We formulate our problem as an optimization problem, show that an optimal solution exists and use generative adversarial networks (GAN) to create perturbations. We further show that our method provides privacy guarantees under the Pufferfish framework, an elegant generalization of the differential privacy that allows for the modeling of prior knowledge on data and correlations. Through experiments, we show that our method outperforms conventional methods in effectively hiding the sensitive attributes while guaranteeing high performance in the target application, for both property inference and training purposes. Finally, we demonstrate through further experiments that once our model learns a privacy-preserving task, such as hiding subjects’ identity, on a group of individuals, it can perform the same task on a separate group with minimal performance drops.