A Target-Agnostic Attack on Deep Models: Exploiting Security Vulnerabilities of Transfer Learning

A Target-Agnostic Attack on Deep Models: Exploiting Security Vulnerabilities of Transfer Learning
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
ArXiv
影响因子:
--
通讯作者:
Shahbaz Rezaei;Xin Liu
Shahbaz Rezaei;Xin Liu
中科院分区:
其他
文献类型:
--
作者:
Shahbaz Rezaei;Xin Liu

文献摘要

被引文献

相似文献

由于训练数据不足以及从头开始训练深度神经网络的计算成本高,迁移学习已被广泛用于许多基于深度神经网络的应用中。一种常用的迁移学习方法是从预先训练好的模型中提取一部分,在最后添加几层,然后用一个小数据集重新训练新的层。这种方法虽然高效且广泛使用,但存在安全漏洞,因为迁移学习中使用的预训练模型通常是公开的,包括潜在的攻击者。在本文中,我们证明了除了预先训练的模型之外,没有任何额外的知识,攻击者可以发起有效和高效的蛮力攻击,可以制作输入的实例,以高置信度触发每个目标类。我们假设攻击者无法访问任何特定于目标的信息,包括来自目标类的样本,重新训练的模型以及Softmax分配给每个类的概率,从而使攻击目标不可知。据我们所知,这些假设使得所有以前的攻击模型都不适用。为了评估所提出的攻击,我们进行了一系列的人脸识别和语音识别任务的实验,并显示了攻击的有效性。我们的工作揭示了Softmax层在迁移学习设置中使用时的基本安全弱点。
Due to insufficient training data and the high computational cost to train a deep neural network from scratch, transfer learning has been extensively used in many deep-neural-network-based applications. A commonly used transfer learning approach involves taking a part of a pre-trained model, adding a few layers at the end, and re-training the new layers with a small dataset. This approach, while efficient and widely used, imposes a security vulnerability because the pre-trained model used in transfer learning is usually publicly available, including to potential attackers. In this paper, we show that without any additional knowledge other than the pre-trained model, an attacker can launch an effective and efficient brute force attack that can craft instances of input to trigger each target class with high confidence. We assume that the attacker has no access to any target-specific information, including samples from target classes, re-trained model, and probabilities assigned by Softmax to each class, and thus making the attack target-agnostic. These assumptions render all previous attack models inapplicable, to the best of our knowledge. To evaluate the proposed attack, we perform a set of experiments on face recognition and speech recognition tasks and show the effectiveness of the attack. Our work reveals a fundamental security weakness of the Softmax layer when used in transfer learning settings.