Gaussian Mixture Models for Classification and Hypothesis Tests Under Differential Privacy

Gaussian Mixture Models for Classification and Hypothesis Tests Under Differential Privacy
复制标题

差分隐私下分类和假设检验的高斯混合模型

DOI:
10.1007/978-3-319-61176-1_7
复制
发表时间:
2017
影响因子:
2.4
通讯作者:
Ali Inan
Ali Inan
中科院分区:
医学4区
文献类型:
--
作者:
Xiaosu Tong;B. Xi;Murat Kantarcioglu;Ali Inan

文献摘要

被引文献

相似文献

许多统计模型都是使用非常基本的统计数据构建的:均值向量、方差和协方差。高斯混合模型就是这样的模型。当数据集包含敏感信息且不能直接发布给用户时,可以基于添加噪声的查询响应轻松构建此类模型。尽管如此,这些模型为用户提供了初步的结果。虽然查询的基本统计信息满足差分隐私保证,但使用这些统计信息构建的复杂模型可能不满足差分隐私保证。然而,如何查询数据库以及如何进一步利用查询结果取决于用户。在本文中,我们的目标是了解差分隐私机制对高斯混合模型的影响。我们的方法包括在差分隐私保护下从数据库中查询基本统计数据,并使用噪声添加响应来构建分类器并进行假设检验。我们发现加入拉普拉斯噪声可能对模型输出有不可忽略的影响。例如,加入噪声后的方差-协方差矩阵不再是正定的。我们提出了一种启发式算法来修复添加方差-协方差矩阵的噪声。然后,我们通过模拟数据和实际数据的实验,使用添加噪声的响应来检查分类误差,并证明在哪些条件下可以减少添加噪声的影响。我们计算了方差相等的单样本z检验、单样本t检验和两样本t检验在差分隐私下的确切类型I和类型II误差。然后,我们显示在何种条件下,假设检验返回可靠的结果给出不同的私有均值,方差和协方差。
Many statistical models are constructed using very basic statistics: mean vectors, variances, and covariances. Gaussian mixture models are such models. When a data set contains sensitive information and cannot be directly released to users, such models can be easily constructed based on noise added query responses. The models nonetheless provide preliminary results to users. Although the queried basic statistics meet the differential privacy guarantee, the complex models constructed using these statistics may not meet the differential privacy guarantee. However it is up to the users to decide how to query a database and how to further utilize the queried results. In this article, our goal is to understand the impact of differential privacy mechanism on Gaussian mixture models. Our approach involves querying basic statistics from a database under differential privacy protection, and using the noise added responses to build classifier and perform hypothesis tests. We discover that adding Laplace noises may have a non-negligible effect on model outputs. For example variance-covariance matrix after noise addition is no longer positive definite. We propose a heuristic algorithm to repair the noise added variance-covariance matrix. We then examine the classification error using the noise added responses, through experiments with both simulated data and real life data, and demonstrate under which conditions the impact of the added noises can be reduced. We compute the exact type I and type II errors under differential privacy for one sample z test, one sample t test, and two sample t test with equal variances. We then show under which condition a hypothesis test returns reliable result given differentially private means, variances and covariances.