Beyond Lp clipping: Equalization-based Psychoacoustic Attacks against ASRs

Beyond Lp clipping: Equalization-based Psychoacoustic Attacks against ASRs
复制标题

DOI:
--
复制
发表时间:
2021-10
期刊:
--
影响因子:
--
通讯作者:
H. Abdullah;Muhammad Sajidur Rahman;Christian Peeters;Cassidy Gibson;Washington Garcia;Vincent Bindschaedler;T. Shrimpton;Patrick Traynor
H. Abdullah;Muhammad Sajidur Rahman;Christian Peeters;Cassidy Gibson;Washington Garcia;Vincent Bindschaedler;T. Shrimpton;Patrick Traynor
中科院分区:
其他
文献类型:
--
作者:
H. Abdullah;Muhammad Sajidur Rahman;Christian Peeters;Cassidy Gibson;Washington Garcia;Vincent Bindschaedler;T. Shrimpton;Patrick Traynor

文献摘要

相似文献

自动语音识别(ASR)系统将语音转换为文本,可以分为两大类:传统的和完全端到端的。这两种类型的音频都被证明容易受到敌意音频示例的攻击,这些示例听起来对人耳是良性的,但迫使ASR产生恶意转录。在这些攻击中,只有“心理声学”攻击可以创建具有相对难以察觉的干扰的例子,因为它们利用了人类听觉系统的知识。不幸的是,现有的心理声学攻击只能应用于传统模型,并且对于较新的、完全端到端的ASR来说已经过时。在本文中,我们提出了一种基于均衡的心理声学攻击,它既可以利用传统的ASR,也可以利用完全端到端的ASR。我们成功地展示了我们对包括DeepSpeech和Wav2Letter在内的真实ASR的攻击。此外,我们使用了一个用户研究来验证我们的方法产生了低的声音失真。具体地说,100名参与者中有80人投票支持我们的所有攻击音频样本,因为它比现有的最先进的攻击噪音更小。通过这一点,我们证明了这两种类型的现有ASR管道可以被利用,以最小的降级来攻击音频质量。
Automatic Speech Recognition (ASR) systems convert speech into text and can be placed into two broad categories: traditional and fully end-to-end. Both types have been shown to be vulnerable to adversarial audio examples that sound benign to the human ear but force the ASR to produce malicious transcriptions. Of these attacks, only the"psychoacoustic"attacks can create examples with relatively imperceptible perturbations, as they leverage the knowledge of the human auditory system. Unfortunately, existing psychoacoustic attacks can only be applied against traditional models, and are obsolete against the newer, fully end-to-end ASRs. In this paper, we propose an equalization-based psychoacoustic attack that can exploit both traditional and fully end-to-end ASRs. We successfully demonstrate our attack against real-world ASRs that include DeepSpeech and Wav2Letter. Moreover, we employ a user study to verify that our method creates low audible distortion. Specifically, 80 of the 100 participants voted in favor of all our attack audio samples as less noisier than the existing state-of-the-art attack. Through this, we demonstrate both types of existing ASR pipelines can be exploited with minimum degradation to attack audio quality.