SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification Systems

SoK: The Faults in our ASRs: An Overview of Attacks against Automatic Speech Recognition and Speaker Identification Systems
复制标题

DOI:
10.1109/sp40001.2021.00014
复制
发表时间:
2020-07
期刊:
2021 IEEE Symposium on Security and Privacy (SP)
影响因子:
--
通讯作者:
H. Abdullah;Kevin Warren;Vincent Bindschaedler;Nicolas Papernot;Patrick Traynor
H. Abdullah;Kevin Warren;Vincent Bindschaedler;Nicolas Papernot;Patrick Traynor
中科院分区:
其他
文献类型:
--
作者:
H. Abdullah;Kevin Warren;Vincent Bindschaedler;Nicolas Papernot;Patrick Traynor

文献摘要

被引文献

相似文献

语音和说话者识别系统被用于各种应用中,从个人助理到电话监控和生物特征认证。这些系统的广泛部署已成为可能的提高精度的神经网络。与其他基于神经网络的系统一样,最近的研究表明,语音和说话人识别系统很容易受到使用操纵输入的攻击。然而,正如我们在本文中所展示的那样,语音和扬声器系统的端到端架构及其输入的性质使得对它们的攻击和防御与图像空间中的攻击和防御有很大不同。我们首先通过系统化该领域的现有研究来证明这一点,并提供一个分类法,通过该分类法,社区可以评估未来的工作。然后,我们通过实验证明,对这些模型的攻击几乎普遍无法转移。在这样做的时候,我们认为,需要大量的额外工作,以提供足够的缓解在这个空间。
Speech and speaker recognition systems are employed in a variety of applications, from personal assistants to telephony surveillance and biometric authentication. The wide deployment of these systems has been made possible by the improved accuracy in neural networks. Like other systems based on neural networks, recent research has demonstrated that speech and speaker recognition systems are vulnerable to attacks using manipulated inputs. However, as we demonstrate in this paper, the end-to-end architecture of speech and speaker systems and the nature of their inputs make attacks and defenses against them substantially different than those in the image space. We demonstrate this first by systematizing existing research in this space and providing a taxonomy through which the community can evaluate future work. We then demonstrate experimentally that attacks against these models almost universally fail to transfer. In so doing, we argue that substantial additional work is required to provide adequate mitigations in this space.