EFFECTIVELY USING PUBLIC DATA IN PRIVACY PRE - SERVING M ACHINE LEARNING
EFFECTIVELY USING PUBLIC DATA IN PRIVACY PRE - SERVING M ACHINE LEARNING
复制标题
在隐私保护中有效使用公共数据 - 服务机器学习
DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
Irfan Khan
中科院分区:
文献类型:
--
作者:
Preeti;Irfan Khan
A key challenge towards differentially private machine learning is balancing the trade-off between privacy and utility. A recent line of work has demonstrated that leveraging public data samples can enhance the utility of DP-trained models (for the same privacy guarantees). In this work, we show that public data can be used to improve utility in DP models significantly more than shown in recent works. Towards this end, we introduce a modified DP-SGD algorithm that leverages public data during its training process. Our technique uses public data in two complementary ways: (1) it uses generative models trained on public data to produce synthetic data that is effectively embedded in multiple steps of the training pipeline; (2) it uses a new gradient clipping mechanism (required for achieving differential privacy) which changes the origin of gradient vectors using information inferred from available public and generated data from generative models. Our experimental results demonstrate the effectiveness of our approach in improving the state-of-the-art in differentially private machine learning across multiple datasets, network architectures, and application domains. Notably, we achieve a 75.1% accuracy on CIFAR10 when using only 2, 000 public images; this is significantly higher than the state-of-the-art which is 68.1% for DP-SGD with the privacy budget of ε = 2, δ = 10−5 (given the same number of public data points).