Machine Learning with Differentially Private Labels: Mechanisms and Frameworks

Machine Learning with Differentially Private Labels: Mechanisms and Frameworks
复制标题

DOI:
10.56553/popets-2022-0112
复制
发表时间:
2022-10
期刊:
Proc. Priv. Enhancing Technol.
影响因子:
--
通讯作者:
Xinyu Tang;Milad Nasr;Saeed Mahloujifar;Virat Shejwalkar;Liwei Song;Amir Houmansadr;Prateek
Xinyu Tang;Milad Nasr;Saeed Mahloujifar;Virat Shejwalkar;Liwei Song;Amir Houmansadr;Prateek
中科院分区:
其他
文献类型:
--
作者:
Xinyu Tang;Milad Nasr;Saeed Mahloujifar;Virat Shejwalkar;Liwei Song;Amir Houmansadr;Prateek

文献摘要

被引文献

相似文献

标签差异隐私是对机器学习场景中差异隐私的一种放松,在这种场景中,标签是训练数据中唯一需要保护的敏感信息。例如,想象一项来自大学班级参与者的关于他们的疫苗接种状况的调查。学生的一些属性是公开的,但他们的疫苗接种状态是敏感信息,必须保密。现在,如果我们想训练一个模型来预测一个学生是否只使用他们的公开信息接种了疫苗,我们可以使用label-DP。最近关于标签- dp的研究使用不同的方法向标签添加噪声以获得标签- dp模型。在这项工作中,我们提出了利用无监督学习和半监督学习来训练具有标签dp保证的模型的新技术,使我们能够在获得相同隐私的同时注入更少的噪声,从而实现更好的效用-隐私权衡。我们首先引入了一个框架,该框架从无监督分类器f0和带有噪声标签集Y的数据集D开始,使用f0减少Y中的噪声,然后使用噪声较小的数据集训练新模型f。我们的降噪策略使用模型f0来去除高概率不正确的噪声标签。然后我们使用半监督学习来训练使用剩余标签的模型。我们用多种方法实例化这个框架来获得噪声标签和基本分类器。作为减少噪声的另一种方法,我们探索了使用无监督学习的效果:我们只在多数投票步骤中添加噪声,以便将学习到的聚类与聚类标签相关联(而不是向单个标签添加噪声);降低的灵敏度使我们能减少噪音。我们的实验表明,这些技术可以显著优于先前在标签- dp上的工作。
Label differential privacy is a relaxation of differential privacy for machine learning scenarios where the labels are the only sensitive information that needs to be protected in the training data. For example, imagine a survey from a participant in a university class about their vaccination status. Some attributes of the students are publicly available but their vaccination status is sensitive information and must remain private. Now if we want to train a model that predicts whether a student has received vaccination using only their public information, we can use label-DP. Recent works on label-DP use different ways of adding noise to the labels in order to obtain label-DP models. In this work, we present novel techniques for training models with label-DP guarantees by leveraging unsupervised learning and semi-supervised learning, enabling us to inject less noise while obtaining the same privacy, therefore achieving a better utility-privacy trade-off. We first introduce a framework that starts with an unsupervised classifier f0 and dataset D with noisy label set Y , reduces the noise in Y using f0 , and then trains a new model f using the less noisy dataset. Our noise reduction strategy uses the model f0 to remove the noisy labels that are incorrect with high probability. Then we use semi-supervised learning to train a model using the remaining labels. We instantiate this framework with multiple ways of obtaining the noisy labels and also the base classifier. As an alternative way to reduce the noise, we explore the effect of using unsupervised learning: we only add noise to a majority voting step for associating the learned clusters with a cluster label (as opposed to adding noise to individual labels); the reduced sensitivity enables us to add less noise. Our experiments show that these techniques can significantly outperform the prior works on label-DP.