Practical Adversarial Attacks Against Speaker Recognition Systems

Practical Adversarial Attacks Against Speaker Recognition Systems
复制标题

DOI:
10.1145/3376897.3377856
复制
发表时间:
2020-02
期刊:
Proceedings of the 21st International Workshop on Mobile Computing Systems and Applications
影响因子:
--
通讯作者:
Zhuohang Li;Cong Shi;Yi Xie;Jian Liu;Bo Yuan;Yingying Chen
Zhuohang Li;Cong Shi;Yi Xie;Jian Liu;Bo Yuan;Yingying Chen
中科院分区:
其他
文献类型:
--
作者:
Zhuohang Li;Cong Shi;Yi Xie;Jian Liu;Bo Yuan;Yingying Chen

文献摘要

被引文献

相似文献

与其他基于生物特征的用户识别方法(例如,指纹和虹膜),说话人识别系统可以依靠他们独特的语音生物特征来识别个人,而不需要用户亲自在场。因此,说话人识别系统近年来在诸如远程访问控制、银行服务和刑事调查的各个领域中变得越来越流行。在本文中,我们通过对最先进的基于深度神经网络(DNN)的说话人识别系统X-vector发起实际和系统的对抗攻击来研究这类系统的脆弱性。特别是,通过在原始音频中添加精心制作的不显眼的噪声,我们的攻击可以欺骗说话人识别系统做出错误的预测,甚至迫使音频被识别为任何对手想要的说话人。此外,我们的攻击将估计的房间脉冲响应(RIR)集成到对抗性示例训练过程中,以实现实际的音频对抗性示例,这些示例在物理世界中通过空中播放时仍然有效。使用109 $扬声器的公共数据集的广泛实验表明,我们的攻击的有效性与高攻击成功率的数字攻击(98%$)和实际的空中攻击(50%$)。
Unlike other biometric-based user identification methods (e.g., fingerprint and iris), speaker recognition systems can identify individuals relying on their unique voice biometrics without requiring users to be physically present. Therefore, speaker recognition systems have been becoming increasingly popular recently in various domains, such as remote access control, banking services and criminal investigation. In this paper, we study the vulnerability of this kind of systems by launching a practical and systematic adversarial attack againstX-vector, the state-of-the-art deep neural network (DNN) based speaker recognition system. In particular, by adding a well-crafted inconspicuous noise to the original audio, our attack can fool the speaker recognition system to make false predictions and even force the audio to be recognized as any adversary-desired speaker. Moreover, our attack integrates the estimated room impulse response (RIR) into the adversarial example training process toward practical audio adversarial examples which could remain effective while being played over the air in the physical world. Extensive experiment using a public dataset of $109$ speakers shows the effectiveness of our attack with a high attack success rate for both digital attack ($98%$) and practical over-the-air attack ($50%$).