The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural Networks

The Secret Revealer: Generative Model-Inversion Attacks Against Deep Neural Networks
复制标题

DOI:
10.1109/cvpr42600.2020.00033
复制
发表时间:
2019-11
期刊:
2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Yuheng Zhang;R. Jia;Hengzhi Pei;Wenxiao Wang;Bo Li;D. Song
Yuheng Zhang;R. Jia;Hengzhi Pei;Wenxiao Wang;Bo Li;D. Song
中科院分区:
其他
文献类型:
--
作者:
Yuheng Zhang;R. Jia;Hengzhi Pei;Wenxiao Wang;Bo Li;D. Song

文献摘要

被引文献

相似文献

本文研究了模型反演攻击,其中对模型的访问被滥用来推断有关训练数据的信息。自从~\cite{fredrikson 2014 privacy}首次引入以来,由于训练数据通常包含隐私敏感信息,此类攻击引起了严重关注。到目前为止,成功的模型反演攻击仅在简单模型上得到证明,例如线性回归和逻辑回归。以前尝试反转神经网络,即使是具有简单架构的神经网络,也未能产生令人信服的结果。在这里,我们提出了一种新的攻击方法,称为生成模型反转攻击,它可以以高成功率反转深度神经网络。我们不是从头开始重建私人训练数据,而是利用部分公共信息(可能非常通用)通过生成对抗网络(GAN)学习分布先验,并使用它来指导反演过程。此外,我们从理论上证明了模型的预测能力及其对反转攻击的脆弱性确实是同一枚硬币的两面-高度预测的模型能够在特征和标签之间建立强相关性,这与对手利用什么来发起攻击完全一致。我们广泛的实验表明,所提出的攻击提高了识别精度超过现有的工作约$75\%$重建的人脸图像从国家的最先进的人脸识别分类。我们还表明,差分隐私,在其规范形式,是无济于事的,以抵御我们的攻击。
This paper studies model-inversion attacks, in which the access to a model is abused to infer information about the training data. Since its first introduction by~\cite{fredrikson2014privacy}, such attacks have raised serious concerns given that training data usually contain privacy sensitive information. Thus far, successful model-inversion attacks have only been demonstrated on simple models, such as linear regression and logistic regression. Previous attempts to invert neural networks, even the ones with simple architectures, have failed to produce convincing results. Here we present a novel attack method, termed the \emph{generative model-inversion attack}, which can invert deep neural networks with high success rates. Rather than reconstructing private training data from scratch, we leverage partial public information, which can be very generic, to learn a distributional prior via generative adversarial networks (GANs) and use it to guide the inversion process. Moreover, we theoretically prove that a model's predictive power and its vulnerability to inversion attacks are indeed two sides of the same coin---highly predictive models are able to establish a strong correlation between features and labels, which coincides exactly with what an adversary exploits to mount the attacks. Our extensive experiments demonstrate that the proposed attack improves identification accuracy over the existing work by about $75\%$ for reconstructing face images from a state-of-the-art face recognition classifier. We also show that differential privacy, in its canonical form, is of little avail to defend against our attacks.