Life after Speech Recognition: Fuzzing Semantic Misinterpretation for Voice Assistant Applications

Life after Speech Recognition: Fuzzing Semantic Misinterpretation for Voice Assistant Applications
复制标题

DOI:
10.14722/ndss.2019.23525
复制
发表时间:
2019
期刊:
Proceedings 2019 Network and Distributed System Security Symposium
影响因子:
--
通讯作者:
Yangyong Zhang;Lei Xu;Abner Mendoza;Guangliang Yang;Phakpoom Chinprutthiwong;G. Gu
Yangyong Zhang;Lei Xu;Abner Mendoza;Guangliang Yang;Phakpoom Chinprutthiwong;G. Gu
中科院分区:
其他
文献类型:
--
作者:
Yangyong Zhang;Lei Xu;Abner Mendoza;Guangliang Yang;Phakpoom Chinprutthiwong;G. Gu

文献摘要

相似文献

—诸如亚马逊Alexa和谷歌助手等流行的语音助手(VA)服务如今正在迅速将其平台应用化,以提供更灵活多样的语音控制服务体验。然而,语音助手设备的广泛部署以及第三方应用数量的不断增加引发了安全和隐私方面的担忧。虽然之前诸如隐蔽语音攻击等研究大多考察了语音助手服务默认的自动语音识别(ASR)组件的问题,但我们的工作分析和评估了自动语音识别之后的后续组件,即自然语言理解(NLU)的安全性,该组件在自动语音识别的语音到文本处理之后进行语义解释(即文本到意图)。特别是,我们专注于用于为第三方语音助手应用(或vApps)定制机器理解的自然语言理解的意图分类器。我们发现,当攻击者巧妙地利用一些常见的口语错误时,意图分类器的不当语义解释所导致的语义不一致会为破坏vApp处理的完整性创造机会。在本文中,我们设计了首个语言模型引导的模糊测试工具,名为LipFuzzer,用于评估意图分类器的安全性,并基于vApps的语音命令模板系统地发现潜在的容易被误解的口语错误。为了引导模糊测试,我们借助统计关系学习(SRL)和新兴的自然语言处理(NLP)技术构建了对抗性语言模型。在评估中,我们成功验证了LipFuzzer的有效性和准确性。我们还使用LipFuzzer对亚马逊Alexa和谷歌助手的vApp平台进行了评估。我们已经确定,现实世界中的很大一部分……
—Popular Voice Assistant (VA) services such as Amazon Alexa and Google Assistant are now rapidly appifying their platforms to allow more flexible and diverse voice-controlled service experience. However, the ubiquitous deployment of VA devices and the increasing number of third-party applications have raised security and privacy concerns. While previous works such as hidden voice attacks mostly examine the problems of VA services’ default Automatic Speech Recognition (ASR) component, our work analyzes and evaluates the security of the succeeding component after ASR, i.e., Natural Language Understanding (NLU), which performs semantic interpretation (i.e., text-to-intent) after ASR’s acoustic-to-text processing. In particular, we focus on NLU’s Intent Classifier which is used in customizing machine understanding for third-party VA Applications (or vApps). We find that the semantic inconsistency caused by the improper semantic interpretation of an Intent Classifier can create the opportunity of breaching the integrity of vApp processing when attackers delicately leverage some common spoken errors. In this paper, we design the first linguistic-model-guided fuzzing tool, named LipFuzzer, to assess the security of Intent Classifier and systematically discover potential misinterpretation-prone spoken errors based on vApps’ voice command templates. To guide the fuzzing, we construct adversarial linguistic models with the help of Statistical Relational Learning (SRL) and emerging Natural Language Processing (NLP) techniques. In evaluation, we have successfully verified the effectiveness and accuracy of LipFuzzer. We also use LipFuzzer to evaluate both Amazon Alexa and Google Assistant vApp platforms. We have identified that a large portion of real-world