Semi-black-box Attacks Against Speech Recognition Systems Using Adversarial Samples

Semi-black-box Attacks Against Speech Recognition Systems Using Adversarial Samples
复制标题

DOI:
10.1109/dyspan.2019.8935789
复制
发表时间:
2019-11
期刊:
2019 IEEE International Symposium on Dynamic Spectrum Access Networks (DySPAN)
影响因子:
--
通讯作者:
Yi Wu;Jian Liu;Yingying Chen;Jerry Q. Cheng
Yi Wu;Jian Liu;Yingying Chen;Jerry Q. Cheng
中科院分区:
其他
文献类型:
--
作者:
Yi Wu;Jian Liu;Yingying Chen;Jerry Q. Cheng

文献摘要

相似文献

近年来,随着自动语音识别(ASR)系统被集成到我们周围的各种设备中,它们的安全漏洞已成为公众日益关注的问题。现有研究表明,作为ASR系统计算核心的深度神经网络(DNN)容易受到故意设计的对抗性攻击。基于梯度下降算法,现有的研究已经成功地产生了敌对样本,可以干扰ASR系统,并产生敌对预期的转录文本设计的对手。这些研究大多模拟白盒攻击,需要了解目标ASR系统中的所有组件。在这项工作中,我们提出了第一个半黑盒攻击ASR系统- Kaldi。我们只需要Kaldi的部分信息,不需要DNN的信息,就可以基于梯度无关遗传算法将恶意命令嵌入到单个音频芯片中。制作的音频片段可以被Kaldi识别为嵌入的恶意命令,同时人类无法察觉。实验表明,我们的攻击可以实现高攻击成功率,对三种类型的音频片段(流行音乐,纯音乐和人类命令)进行不明显的扰动,而不需要底层DNN模型参数和架构。
As automatic speech recognition (ASR) systems have been integrated into a diverse set of devices around us in recent years, security vulnerabilities of them have become an increasing concern for the public. Existing studies have demonstrated that deep neural networks (DNNs), acting as the computation core of ASR systems, is vulnerable to deliberately designed adversarial attacks. Based on the gradient descent algorithm, existing studies have successfully generated adversarial samples which can disturb ASR systems and produce adversary-expected transcript texts designed by adversaries. Most of these research simulated white-box attacks which require knowledge of all the components in the targeted ASR systems. In this work, we propose the first semi-black-box attack against the ASR system - Kaldi. Requiring only partial information from Kaldi and none from DNN, we can embed malicious commands into a single audio chip based on the gradient-independent genetic algorithm. The crafted audio clip could be recognized as the embedded malicious commands by Kaldi and unnoticeable to humans in the meanwhile. Experiments show that our attack can achieve high attack success rate with unnoticeable perturbations to three types of audio clips (pop music, pure music, and human command) without the need of the underlying DNN model parameters and architecture.