Locally Differentially Private and Fair Key-Value Aggregation

Locally Differentially Private and Fair Key-Value Aggregation
复制标题

DOI:
10.1145/3628797.3628807
复制
发表时间:
2023-12
期刊:
Proceedings of the 12th International Symposium on Information and Communication Technology
影响因子:
--
通讯作者:
Yukun Dong;Zhengwu Lu;Rui Zhang
Yukun Dong;Zhengwu Lu;Rui Zhang
中科院分区:
其他
文献类型:
--
作者:
Yukun Dong;Zhengwu Lu;Rui Zhang

文献摘要

相似文献

在大数据时代,从庞大的数据集中提取有意义的见解,同时维护个人隐私的能力已经成为一个越来越复杂的挑战。近年来,已经见证了各种本地差异隐私数据聚合方案的发展,这些方案允许不受信任的数据收集者从用户数据中获得有意义的统计数据,同时为个人用户保持强有力的隐私保证。作为NoSQL数据库中的基本数据类型,键值数据有两个重要的统计数据,每个键的频率和相应的平均值。当前的局部差分私钥-值聚合方案主要依赖于用于均值估计的均匀采样,即,从每个用户的键值集中随机选择单个键值对。然而,这种方法导致频繁键的高平均估计精度和不频繁键的低精度。为了解决这个问题,本文提出了自适应的设计和评估,一种新的本地差分私有和公平的键值聚合方案,可以提供统一的高平均估计精度在不同的键。在第一阶段,我们利用隐私预算的一部分来估计每个密钥的频率。随后,基于在第一阶段中估计的键频率,我们采用非均匀随机采样进行均值估计,这使得与低频键相关联的值的概率更高。全面的理论分析和仿真研究证实了自适应比以前的解决方案的优越性。
In the era of Big Data, the ability to extract meaningful insights from vast datasets while maintaining individual privacy has become an increasingly complex challenge. Recent years have witnessed the development of various locally differentially private data aggregation schemes which allow an untrusted data collector to derive meaningful statistics from user data while maintaining strong privacy guarantee for individual users. As a fundamental data type in NoSQL databases, key-value data has two important statistics of interest, the frequency of each key and the corresponding mean value. Current locally differentially private key-value aggregation schemes primarily rely on uniform sampling for mean estimation, i.e., a single key-value pair is selected randomly from each user’s key-value set. This approach, however, results in high mean estimation accuracy for frequent keys and low accuracy for infrequent ones. To tackle this problem, this paper presents the design and evaluation of Adaptive, a novel locally differentially private and fair key-value aggregation scheme that can deliver uniformly high mean estimation accuracy across different keys. In the first phase, we utilize a portion of the privacy budget to estimate the frequency of each key. Subsequently, based on the key frequencies estimated in the first phase, we employ non-uniform random sampling for mean estimation, which enables higher probability sampling of values associated with low-frequency keys. Comprehensive theoretical analysis and simulation studies confirm the superiority of Adaptive over previous solutions.