A Neural Database for Differentially Private Spatial Range Queries

A Neural Database for Differentially Private Spatial Range Queries
复制标题

DOI:
10.14778/3510397.3510404
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Sepanta Zeighami;Ritesh Ahuja;G. Ghinita;C. Shahabi
Sepanta Zeighami;Ritesh Ahuja;G. Ghinita;C. Shahabi
中科院分区:
其他
文献类型:
--
作者:
Sepanta Zeighami;Ritesh Ahuja;G. Ghinita;C. Shahabi

文献摘要

相似文献

移动应用程序和基于位置的服务生成大量位置数据。来自此类数据集的位置密度信息受益于交通优化,上下文意见通知和公共卫生的研究(例如疾病传播)。为了保留个人隐私,必须对位置数据进行消毒,这通常是使用差异隐私(DP)进行的。现有方法将数据域分配到垃圾箱中,向每个垃圾箱添加噪声并发布数据的嘈杂直方图。但是,这种简单的建模选择无法准确捕获空间数据集中有用的密度信息并产生差的精度。我们提出了一种基于机器学习的方法,用于回答具有DP保证的位置数据的范围计数查询。我们专注于通过学习来反击困扰现有方法(即噪声和均匀性误差)的错误来源,并且我们设计了一个神经数据库系统,该系统模拟空间数据,以便保留了密度特征,即使添加了DP符合DP的噪声也是如此。我们还设计了一个在公共数据之上进行有效系统参数调整的框架,该框架有助于设置重要的系统参数,而无需花费稀缺的隐私预算。具有异质特征的真实数据集的广泛实验结果表明,我们提出的方法显着优于最新技术。
Mobile apps and location-based services generate large amounts of location data. Location density information from such datasets benefits research on traffic optimization, context-aware notifications and public health (e.g., disease spread). To preserve individual privacy, one must sanitize location data, which is commonly done using differential privacy (DP). Existing methods partition the data domain into bins, add noise to each bin and publish a noisy histogram of the data. However, such simplistic modelling choices fall short of accurately capturing the useful density information in spatial datasets and yield poor accuracy. We propose a machine-learning based approach for answering range count queries on location data with DP guarantees. We focus on countering the sources of error that plague existing approaches (i.e., noise and uniformity error) through learning, and we design a neural database system that models spatial data such that density features are preserved, even when DP-compliant noise is added. We also devise a framework for effective system parameter tuning on top of public data, which helps set important system parameters without expending scarce privacy budget. Extensive experimental results on real datasets with heterogeneous characteristics show that our proposed approach significantly outperforms the state of the art.