EFFECTIVELY USING PUBLIC DATA IN PRIVACY PRE - SERVING M ACHINE LEARNING

EFFECTIVELY USING PUBLIC DATA IN PRIVACY PRE - SERVING M ACHINE LEARNING
复制标题

在隐私保护中有效使用公共数据 - 服务机器学习

DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Irfan Khan
Irfan Khan
中科院分区:
--
文献类型:
--
作者:
Preeti;Irfan Khan

文献摘要

相似文献

区别对待私有机器学习的一个关键挑战是平衡隐私和效用之间的权衡。最近的一系列工作表明,利用公共数据样本可以增强DP训练的模型的实用性(对于相同的隐私保证)。在这项工作中,我们证明了公开数据可以显著地提高DP模型中的效用,而不是最近的工作中所显示的。为此,我们引入了一种改进的DP-SGD算法,该算法在其训练过程中利用了公共数据。我们的技术以两种互补的方式使用公共数据:(1)它使用在公共数据上训练的生成模型来产生合成数据,该合成数据有效地嵌入到训练管道的多个步骤中;(2)它使用新的梯度裁剪机制(实现差异化隐私所必需的),该机制使用从可用的公共数据推断的信息和从生成模型生成的信息来改变梯度向量的来源。我们的实验结果证明了我们的方法在提高跨多个数据集、网络体系结构和应用领域的差异私有机器学习方面的有效性。值得注意的是,在仅使用2000个公共图像的情况下,我们在CIFAR10上达到了75.1%的准确率;这远远高于最先进的DP-SGD,在隐私预算为ε=2,δ=10−5(给定相同数量的公共数据点)的情况下,DP-SGD的准确率为68.1%。
A key challenge towards differentially private machine learning is balancing the trade-off between privacy and utility. A recent line of work has demonstrated that leveraging public data samples can enhance the utility of DP-trained models (for the same privacy guarantees). In this work, we show that public data can be used to improve utility in DP models significantly more than shown in recent works. Towards this end, we introduce a modified DP-SGD algorithm that leverages public data during its training process. Our technique uses public data in two complementary ways: (1) it uses generative models trained on public data to produce synthetic data that is effectively embedded in multiple steps of the training pipeline; (2) it uses a new gradient clipping mechanism (required for achieving differential privacy) which changes the origin of gradient vectors using information inferred from available public and generated data from generative models. Our experimental results demonstrate the effectiveness of our approach in improving the state-of-the-art in differentially private machine learning across multiple datasets, network architectures, and application domains. Notably, we achieve a 75.1% accuracy on CIFAR10 when using only 2, 000 public images; this is significantly higher than the state-of-the-art which is 68.1% for DP-SGD with the privacy budget of ε = 2, δ = 10−5 (given the same number of public data points).