Local Information Privacy and Its Application to Privacy-Preserving Data Aggregation

Local Information Privacy and Its Application to Privacy-Preserving Data Aggregation
复制标题

DOI:
10.1109/tdsc.2020.3041733
复制
发表时间:
2022-05-01
影响因子:
7.3
通讯作者:
Tandon, Ravi
Tandon, Ravi
中科院分区:
计算机科学2区
文献类型:
--
作者:
Jiang, Bo;Li, Ming;Tandon, Ravi

文献摘要

被引文献

相似文献

在这篇文章中,我们提出了本地信息隐私(LIP),并设计了基于LIP的统计聚合机制,同时保护用户的隐私,而不依赖于可信的第三方。LIP引入了上下文感知的概念,它可以被看作是利用数据的先验(包括私有化和后处理)来提高数据的效用。我们提出了一个优化框架,以最大限度地减少数据聚合的均方误差,同时保护每个用户的输入数据或相关的潜在变量的隐私,通过满足LIP约束。然后,我们研究了在不同的情况下,考虑先验的不确定性和相关性的潜变量的最优机制。在这篇文章中,包括随机响应(RR),一元编码(UE),和本地哈希(LH)的三种类型的机制进行了研究,我们推导出封闭形式的解决方案的最佳扰动参数是先验相关的。我们比较了基于LIP的机制和基于LDP的机制,并从理论上证明了前者实现了增强的效用。然后,我们研究两个应用程序:(加权)求和和直方图估计,并显示如何提出的机制可以应用到每个应用程序。最后,我们验证了我们的分析模拟使用合成和真实世界的数据。结果显示了不同的先验分布,相关性和输入域大小对数据效用的影响。结果还表明,我们的LIP为基础的机制提供了更好的实用性,隐私权衡比基于LDP的。
In this article, we propose local information privacy (LIP), and design LIP based mechanisms for statistical aggregation while protecting users' privacy without relying on a trusted third party. The concept of context-awareness is incorporated in LIP, which can be viewed as exploiting of data prior (both in privatizing and post-processing) to enhance data utility. We present an optimization framework to minimize the mean square error of data aggregation while protecting the privacy of each user's input data or a correlated latent variable by satisfying LIP constraints. Then, we study optimal mechanisms under different scenarios considering the prior uncertainty and correlation with a latent variable. Three types of mechanisms are studied in this article, including randomized response (RR), unary encoding (UE), and local hashing (LH), and we derive closed-form solutions for the optimal perturbation parameters that are prior-dependent. We compare LIP-based mechanisms with those based on LDP, and theoretically show that the former achieve enhanced utility. We then study two applications: (weighted) summation and histogram estimation, and show how proposed mechanisms can be applied to each application. Finally, we validate our analysis by simulations using both synthetic and real-world data. Results show the impact on data utility by different prior distributions, correlations, and input domain sizes. Results also show that our LIP-based mechanisms provide better utility-privacy tradeoffs than LDP-based ones.