Robust Detection of Machine-induced Audio Attacks in Intelligent Audio Systems with Microphone Array

Robust Detection of Machine-induced Audio Attacks in Intelligent Audio Systems with Microphone Array
复制标题

DOI:
10.1145/3460120.3484755
复制
发表时间:
2021-11
期刊:
Proceedings of the 2021 ACM SIGSAC Conference on Computer and Communications Security
影响因子:
--
通讯作者:
Zhuohang Li;Cong Shi;Tianfang Zhang;Yi Xie;Jian Liu;Bo Yuan;Yingying Chen
Zhuohang Li;Cong Shi;Tianfang Zhang;Yi Xie;Jian Liu;Bo Yuan;Yingying Chen
中科院分区:
其他
文献类型:
--
作者:
Zhuohang Li;Cong Shi;Tianfang Zhang;Yi Xie;Jian Liu;Bo Yuan;Yingying Chen

文献摘要

相似文献

随着近年来智能音频系统的普及,其漏洞已成为公众日益关注的问题。现有的研究已经设计了一系列机器诱导的音频攻击,如重放攻击、合成攻击、隐藏语音命令、听不见的攻击和音频对抗示例,这些攻击可能会使用户面临严重的安全和隐私威胁。为了防御这些攻击,现有的努力一直是单独对待它们。虽然它们在某些情况下产生了相当好的性能,但在实践中,它们很难组合成一个多功能一体的解决方案来部署在音频系统上。此外,现代智能音频设备,如Amazon Echo和Apple HomePod,通常配备麦克风阵列,用于远场语音识别和降噪。现有的防御策略主要集中在单通道和双通道音频上,而很少有研究探索使用多通道麦克风阵列来防御特定类型的音频攻击。针对目前对防范各种音频攻击缺乏系统研究的现状,以及多通道音频的潜在优势,本文提出了一种基于多通道麦克风阵列的智能音频系统检测机器诱导音频攻击的整体解决方案。具体来说,我们利用多声道音频的幅度和相位谱图来提取空间信息,并利用深度学习模型来检测人类语音和回放机器生成的对抗性音频之间的根本差异。此外,我们采用了一个无监督的域自适应训练框架,以进一步提高模型的泛化能力,在新的声学环境。在各种设置下对公共多通道重放攻击数据集和自我收集的涉及5种高级音频攻击的多通道音频攻击数据集进行评估。实验结果表明,该方法在检测各种机器诱导攻击时,可以达到低至6.6%的等错误率(EER)。即使在新的声学环境中,我们的方法仍然可以实现低至8.8%的EER。
With the popularity of intelligent audio systems in recent years, their vulnerabilities have become an increasing public concern. Existing studies have designed a set of machine-induced audio attacks, such as replay attacks, synthesis attacks, hidden voice commands, inaudible attacks, and audio adversarial examples, which could expose users to serious security and privacy threats. To defend against these attacks, existing efforts have been treating them individually. While they have yielded reasonably good performance in certain cases, they can hardly be combined into an all-in-one solution to be deployed on the audio systems in practice. Additionally, modern intelligent audio devices, such as Amazon Echo and Apple HomePod, usually come equipped with microphone arrays for far-field voice recognition and noise reduction. Existing defense strategies have been focusing on single- and dual-channel audio, while only few studies have explored using multi-channel microphone array for defending specific types of audio attack. Motivated by the lack of systematic research on defending miscellaneous audio attacks and the potential benefits of multi-channel audio, this paper builds a holistic solution for detecting machine-induced audio attacks leveraging multi-channel microphone arrays on modern intelligent audio systems. Specifically, we utilize magnitude and phase spectrograms of multi-channel audio to extract spatial information and leverage a deep learning model to detect the fundamental difference between human speech and adversarial audio generated by the playback machines. Moreover, we adopt an unsupervised domain adaptation training framework to further improve the model's generalizability in new acoustic environments. Evaluation is conducted under various settings on a public multi-channel replay attack dataset and a self-collected multi-channel audio attack dataset involving 5 types of advanced audio attacks. The results show that our method can achieve an equal error rate (EER) as low as 6.6% in detecting a variety of machine-induced attacks. Even in new acoustic environments, our method can still achieve an EER as low as 8.8%.