课题基金 / 基金详情

Privacy Preserving synthesized data releasing via generative adversarial networks

Privacy Preserving synthesized data releasing via generative adversarial networks
通过生成对抗网络发布的隐私保护合成数据
批准号:
RGPIN-2019-06119
负责人:
Wang, Ke
金额:
$3.5万
依托单位:
依托单位国家:
加拿大
项目类别:
Discovery Grants Program - Individual
财政年份:
2022
资助国家:
加拿大
项目状态:
已结题
起止时间:
2022-01-01 至 2023-12-31

项目摘要

项目成果

Wang, Ke的其他基金

相似基金

相关文献

中文摘要
翻译
最近的Facebook隐私丑闻涉及伦敦一家数据挖掘公司滥用Facebook数千万用户的信息,再次凸显了收集、存储和使用敏感数据进行数据分析的隐私问题。传统的数据清理技术通过屏蔽真实数据中的敏感信息并发布屏蔽版本来解决隐私问题。这种方法对什么是敏感信息以及如何使用这些数据做了一定的假设。通常这些信息是不可用的,因此过度的数据清理是必要的,这破坏了潜在分析的数据效用。本研究的目的是探讨通过保留真实数据的分布特征来发布真实数据生成的合成数据的替代方案,而不是发布实际的个人记录。关键是如何保持分布特征,以及如何确保这样做不会泄露有关个人的敏感信息。最近机器学习和深度神经网络中生成对抗网络(GANs)的发展为解决上述问题开辟了新的可能性。GANs是两个神经网络在零和博弈框架下相互竞争的系统。生成网络学习从潜在空间映射到感兴趣的特定数据分布,而判别网络区分真实数据分布的实例和生成网络产生的候选实例。两种网络都改进了它们的方法,直到合成的实例与真实的实例无法区分,即保留了真实数据的分布特征。为了解决隐私问题,以前的工作主要来自图像生成,在gan训练的随机梯度下降过程中加入随机噪声来扰动梯度。由于扰动梯度对解的收敛速度和效用有不利影响,因此由于效用较差,仅对弱隐私设置进行了评估。拟议的研究将研究添加噪声的替代方法,这些方法可以更好地保护隐私和实用性,并评估数据在广泛领域的实用性。我们特别感兴趣的一个应用是向研究人员发布医疗和保健数据,这要归功于我们在该领域的真实数据和专业知识。本研究的意义在于,数据持有人不必担心数据隐私,因为没有真实的数据被发布,发布的数据有很强的隐私保障;另一方面,数据分析师将得到几乎相同的结果,就好像分析了真实数据一样。这项工作将有助于隐私保护的实践和鼓励数据共享,以实现数据分析的利益。
英文摘要
Privacy Preserving synthesized data releasing via generative adversarial networks The recent Facebook privacy scandal involving a London-based data-mining firm on misusing Facebook information of tens of millions of users highlights again the privacy concern over collecting, storing, and using sensitive data for data analysis. The traditional data sanitization technique addresses privacy concerns by masking sensitive information in true data and releasing the masked version. This approach makes certain assumptions on what is sensitive information and how the data will be used. Often these information are not available, so excessive data sanitization is necessary, which destroys data utility for potential analyses. The objective of this proposed research is to investigate the alternative of releasing synthesized data generated from true data by preserving the distributional characteristics of true data, instead of releasing actual individuals' records. The key is how to preserve distributional characteristics and how to ensure that doing so does not disclose sensitive information about individuals. The recent development of Generative Adversarial Networks (GANs) in machine learning and deep neural networks opens up new possibilities to address the above problem. GANs are a system of two neural networks contesting with each other in a zero-sum game framework. The generative network learns to map from a latent space to a particular data distribution of interest, while the discriminative network discriminates between instances from the true data distribution and candidates produced by the generative network. Both networks improve their methods until the synthesized instances are indistinguishable from the genuine ones, i.e., preserve the distributional characteristics of true data. To address privacy concerns, previous works, mainly from image generation, added random noises to perturb the gradient during stochastic gradient descent in the training of GANs. Since the perturbed gradient adversely affects the convergence rate and the utility of solutions, only weak privacy settings were evaluated because of poor utility. The proposed research will investigate alternatives ways of adding noises that could better preserve both privacy and utility, and evaluate data utility in a broad range of domains. One application of special interests to us is releasing medical and healthcare data to researchers, thanks to our access to true data and expertise in this domain. The significance of this research is that the data holder does not have to be concerned with data privacy because no true data is released and the released data meets a strong privacy guarantee; on the other hand, the data analyst will get nearly the same result as if true data were analyzed. This work will contribute to the practice of privacy preservation and the encouragement of data sharing for the benefits of data analysis.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Privacy Preserving synthesized data releasing via generative adversarial networks
  • 批准号:
    RGPIN-2019-06119
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.5万
  • 财政年份:
    2021
  • 负责人:
    Wang, Ke
  • 依托单位:
Privacy Preserving synthesized data releasing via generative adversarial networks
  • 批准号:
    RGPIN-2019-06119
  • 项目类别:
    Discovery Grants Program - Individual
  • 资助金额:
    $3.5万
  • 财政年份:
    2020
  • 负责人:
    Wang, Ke
  • 依托单位:
Privacy Preserving synthesized data releasing via generative adversarial networks
  • 批准号:
    RGPAS-2019-00081
  • 项目类别:
    Discovery Grants Program - Accelerator Supplements
  • 资助金额:
    $5.83万
  • 财政年份:
    2020
  • 负责人:
    Wang, Ke
  • 依托单位:
Privacy Preserving synthesized data releasing via generative adversarial networks
  • 批准号:
    RGPAS-2019-00081
  • 项目类别:
    Discovery Grants Program - Accelerator Supplements
  • 资助金额:
    $2.91万
  • 财政年份:
    2019
  • 负责人:
    Wang, Ke
  • 依托单位:
海外基金