A Targeted Privacy-Preserving Data Publishing Method Based on Bayesian Network

A Targeted Privacy-Preserving Data Publishing Method Based on Bayesian Network
复制标题

一种基于贝叶斯网络的定向隐私保护数据发布方法

DOI:
10.1109/access.2022.3201641
复制
发表时间:
2022
期刊:
影响因子:
3.9
通讯作者:
Junzhong Miao
Junzhong Miao
中科院分区:
计算机科学3区
文献类型:
--
作者:
Zhigang Zhou;Yu Wang;Xiao Yu;Junzhong Miao

文献摘要

参考文献

相似文献

隐私保护数据发布(PPDP)是数据驱动的AI技术(如数据挖掘、机器学习、深度学习等)的必要前提。安全、合法地从数据中提取知识。在过去的十年里,它作为一个热门话题得到了应有的研究和探索。然而,现有的隐私保护机制没有考虑到以下三个方面:防止后台攻击、最大化数据可用性和抵抗敏感信息挖掘。在这项工作中,我们提出了一种新颖的隐私保护数据发布框架,通过发布模拟数据而不是真实数据来保护隐私。探索了利用贝叶斯网络生成与真实数据分布相似的数据。它由两种成分组成。首先将数据发布问题转化为贝叶斯网络的生成过程,相应地将隐私泄露问题转化为一种贝叶斯推理攻击。其次,我们提出了一种重新匿名框架,称为(d,L)注入,该框架灵活地解决了隐私保护强度增加对数据可用性的影响。此外,我们将三种经典的隐私保护策略移植到生成的贝叶斯网络中,并通过来自多个应用领域的三个公共数据集验证了该方法的有效性。
Privacy-preserving data publishing (PPDP) is an essential prerequisite for data-driven AI technologies, (such as data mining, machine learning, deep learning, etc.) to extract knowledge from data safely and legally. It has, as it should be, been studied and explored as a hot topic in the last decade. However, existing privacy protection mechanisms cannot take into account the following three aspects: preventing background attack, maximizing data availability, and resisting sensitive information mining. In this work, we propose a novel privacy-preserving data publishing framework, which protects privacy by releasing simulated data instead of real data. It is explored for generating data similar to the distribution of the real data by using Bayesian network. It consists of two ingredients. First, we transform the problem of data publication into the generation process of a Bayesian network, and correspondingly, the problem of privacy leakage is transformed into one kind of Bayesian inference attack. Second, we propose a re-anonymity framework, named (d, L)-injection, which flexibly resolves the impact of increased privacy protection strength on data availability. In addition, we transplant three classical privacy-preserving strategies to the generated Bayesian network, and demonstrates the effectiveness of the method through three public data sets from multiple application domains.
DOI: 10.1109/infcom.2010.5462174
发表时间: 2010-03
期刊: 2010 Proceedings IEEE INFOCOM
影响因子: --
作者:
Shucheng Yu;Cong Wang-;K. Ren;Wenjing Lou
通讯作者: Shucheng Yu;Cong Wang-;K. Ren;Wenjing Lou
DOI: 10.1109/infcom.2013.6567072
发表时间: 2013-04
期刊: 2013 Proceedings IEEE INFOCOM
影响因子: --
作者:
Zhigang Zhou;Hongli Zhang;Xiaojiang Du;Panpan Li;Xiangzhan Yu
通讯作者: Zhigang Zhou;Hongli Zhang;Xiaojiang Du;Panpan Li;Xiangzhan Yu
DOI: 10.1109/infocom.2019.8737494
发表时间: 2019-04
期刊: IEEE INFOCOM 2019 - IEEE Conference on Computer Communications
影响因子: --
作者:
Liyao Xiang;Jingbo Yang;Baochun Li
通讯作者: Liyao Xiang;Jingbo Yang;Baochun Li
DOI: 10.1007/978-3-030-68799-1_32
发表时间: 2020
期刊: --
影响因子: --
作者:
J. Wu;Gautam Srivastava;Shahab Tayeb;Chun-Wei Lin
通讯作者: J. Wu;Gautam Srivastava;Shahab Tayeb;Chun-Wei Lin
DOI: 10.1109/pst.2013.6596033
发表时间: 2013-07
期刊: 2013 Eleventh Annual Conference on Privacy, Security and Trust
影响因子: --
作者:
Jordi Soria-Comas;J. Domingo-Ferrer
通讯作者: Jordi Soria-Comas;J. Domingo-Ferrer