Differentially Private Empirical Risk Minimization

Differentially Private Empirical Risk Minimization
复制标题

DOI:
10.5555/1953048.2021036
复制
发表时间:
2009-11
期刊:
Journal of machine learning research : JMLR
影响因子:
--
通讯作者:
Kamalika Chaudhuri;C. Monteleoni;A. Sarwate
Kamalika Chaudhuri;C. Monteleoni;A. Sarwate
中科院分区:
其他
文献类型:
--
作者:
Kamalika Chaudhuri;C. Monteleoni;A. Sarwate

文献摘要

被引文献

相似文献

隐私保护机器学习算法对于日益常见的分析个人数据(如医疗或财务记录)的环境至关重要。我们提供了产生通过(正则化的)经验风险最小化(ERM)学习的分类器的隐私保护近似的通用技术。由于DWork等人的原因,这些算法在ε-Differential隐私定义下是私有的。(2006年)。首先,我们应用了Dwork等人的输出摄动思想。(2006),归入环境风险管理分类。然后,我们提出了一种新的隐私保护机器学习算法设计方法--目标摄动法。这种方法需要在优化过分类器之前对目标函数进行扰动。如果损失和正则化满足一定的凸性和可微性准则,我们证明了我们的算法具有保密性,并提供了线性核和非线性核的广义界。我们进一步提出了一种隐私保护技术,用于调整一般机器学习算法中的参数,从而为训练过程提供端到端的隐私保证。我们将这些结果应用于生成正则化Logistic回归和支持向量机的隐私保护类似物。通过评估它们在真实人口统计和基准数据集上的表现,我们获得了令人鼓舞的结果。我们的结果表明,无论是在理论上还是在经验上,客观扰动在管理隐私和学习成绩之间的内在权衡方面都优于以前的最新的输出扰动。
Privacy-preserving machine learning algorithms are crucial for the increasingly common setting in which personal data, such as medical or financial records, are analyzed. We provide general techniques to produce privacy-preserving approximations of classifiers learned via (regularized) empirical risk minimization (ERM). These algorithms are private under the ε-differential privacy definition due to Dwork et al. (2006). First we apply the output perturbation ideas of Dwork et al. (2006), to ERM classification. Then we propose a new method, objective perturbation, for privacy-preserving machine learning algorithm design. This method entails perturbing the objective function before optimizing over classifiers. If the loss and regularizer satisfy certain convexity and differentiability criteria, we prove theoretical results showing that our algorithms preserve privacy, and provide generalization bounds for linear and nonlinear kernels. We further present a privacy-preserving technique for tuning the parameters in general machine learning algorithms, thereby providing end-to-end privacy guarantees for the training process. We apply these results to produce privacy-preserving analogues of regularized logistic regression and support vector machines. We obtain encouraging results from evaluating their performance on real demographic and benchmark data sets. Our results show that both theoretically and empirically, objective perturbation is superior to the previous state-of-the-art, output perturbation, in managing the inherent tradeoff between privacy and learning performance.