Gradual Release of Sensitive Data under Differential Privacy

Gradual Release of Sensitive Data under Differential Privacy
复制标题

差异隐私下敏感数据逐步发布

DOI:
10.29012/jpc.v7i2.649
复制
发表时间:
2015
期刊:
J. Priv. Confidentiality
影响因子:
--
通讯作者:
George Pappas
George Pappas
中科院分区:
--
文献类型:
--
作者:
Fragkiskos Koufogiannis;Shuo Han;George Pappas

文献摘要

被引文献

相似文献

当隐私级别随时间变化时,我们引入了在不同隐私下发布敏感数据的问题。现有的工作假定在敏感数据发布之前,系统设计者将隐私级别确定为固定值。然而,对于某些应用程序,用户可能希望在重新评估隐私问题或需要更好的准确性之后,放宽相同数据的后续发布的隐私级别。具体地说,在给定包含敏感数据的数据库的情况下,我们假设保留$\epsilon_{1}$-差异隐私的响应$y_1$已经发布。然后,隐私级别放宽到$\epsilon_2$,其中$\epsilon_2;\epsilon_1$,我们希望发布更准确的响应$y_2$,而联合响应$(y_1,y_2)$保留$\epsilon_2$-差异隐私。在逐渐释放两个响应$y_1$和$y_2$的方案中,与释放单个响应是$\epsilon_{2}$-差分私有的方案相比,精确度损失了多少?我们的结果表明,存在一种复合机制,它在精度上实现了零损失。我们考虑了私有数据位于$\mathbb{R}^{n}$内且由$\ell_{1}$-范数诱导的邻接关系的情况,重点讨论了近似身份查询的机制。我们证明了在逐步释放的情况下,通过其输出可以用滞后马尔可夫随机过程来描述的机制,也可以达到同样的精度。这种随机过程具有闭合形式的表达式,可以有效地进行采样。我们的结果不仅适用于身份查询,而且还适用于身份查询。为此,我们展示了我们的结果可以应用于几个案例,包括谷歌的RAPPOR项目,敏感数据的交易,以及社交网络中私人数据的受控传输。
We introduce the problem of releasing sensitive data under differential privacy when the privacy level is subject to change over time. Existing work assumes that privacy level is determined by the system designer as a fixed value before sensitive data is released. For certain applications, however, users may wish to relax the privacy level for subsequent releases of the same data after either a re-evaluation of the privacy concerns or the need for better accuracy. Specifically, given a database containing sensitive data, we assume that a response $y_1$ that preserves $\epsilon_{1}$-differential privacy has already been published. Then, the privacy level is relaxed to $\epsilon_2$, with $\epsilon_2 > \epsilon_1$, and we wish to publish a more accurate response $y_2$ while the joint response $(y_1, y_2)$ preserves $\epsilon_2$-differential privacy. How much accuracy is lost in the scenario of gradually releasing two responses $y_1$ and $y_2$ compared to the scenario of releasing a single response that is $\epsilon_{2}$-differentially private? Our results show that there exists a composite mechanism that achieves \textit{no loss} in accuracy. We consider the case in which the private data lies within $\mathbb{R}^{n}$ with an adjacency relation induced by the $\ell_{1}$-norm, and we focus on mechanisms that approximate identity queries. We show that the same accuracy can be achieved in the case of gradual release through a mechanism whose outputs can be described by a \textit{lazy Markov stochastic process}. This stochastic process has a closed form expression and can be efficiently sampled. Our results are applicable beyond identity queries. To this end, we demonstrate that our results can be applied in several cases, including Google's RAPPOR project, trading of sensitive data, and controlled transmission of private data in a social network.