Read Between the Lines: An Empirical Measurement of Sensitive Applications of Voice Personal Assistant Systems

Read Between the Lines: An Empirical Measurement of Sensitive Applications of Voice Personal Assistant Systems
复制标题

DOI:
10.1145/3366423.3380179
复制
发表时间:
2020-04
期刊:
Proceedings of The Web Conference 2020
影响因子:
--
通讯作者:
F. H. Shezan;Hang Hu;Jiamin Wang;Gang Wang;Yuan Tian
F. H. Shezan;Hang Hu;Jiamin Wang;Gang Wang;Yuan Tian
中科院分区:
其他
文献类型:
--
作者:
F. H. Shezan;Hang Hu;Jiamin Wang;Gang Wang;Yuan Tian

文献摘要

被引文献

相似文献

亚马逊Alexa和Google Home等语音个人助理(VPA)系统已被数千万家庭使用。最近的工作证明了针对其语音接口的概念验证攻击,以调用非预期的应用程序或操作。然而,对于VPA系统支持哪种类型的第三方应用程序,以及这些攻击可能导致什么后果,仍然缺乏经验性的了解。在本文中,我们对Amazon Alexa和Google Home的第三方应用程序进行了实证分析,以系统地评估攻击面。一个关键的方法是通过对它接受的敏感语音命令进行分类来表征给定的应用程序。我们开发了一个自然语言处理工具,从两个维度对给定的语音命令进行分类:(1)语音命令是否旨在插入动作或检索信息;(2)命令是否敏感或不敏感。该工具结合了深度神经网络和基于关键字的模型,并使用主动学习来减少手动标记工作。敏感度分类是基于用户研究(N=404),我们测量感知的语音命令的敏感度。地面实况评估表明,我们的工具在两种类型的分类中均达到了95%以上的准确率。我们应用该工具分析了两年(2018-2019年)内的77,957个Amazon Alexa应用程序和4,813个Google Home应用程序(198,199个来自Amazon Alexa的语音命令,13,644个来自Google Home的语音命令)。我们总共识别出19,263个敏感的“动作注入”命令和5,352个敏感的“信息检索”命令。这些命令来自4,596个应用程序(占所有应用程序的5.55%),其中大部分属于“智能家居”类别。虽然敏感应用程序的比例很小,但我们发现,从2018年到2019年,这一比例正随着时间的推移而增加。
Voice Personal Assistant (VPA) systems such as Amazon Alexa and Google Home have been used by tens of millions of households. Recent work demonstrated proof-of-concept attacks against their voice interface to invoke unintended applications or operations. However, there is still a lack of empirical understanding of what type of third-party applications that VPA systems support, and what consequences these attacks may cause. In this paper, we perform an empirical analysis of the third-party applications of Amazon Alexa and Google Home to systematically assess the attack surfaces. A key methodology is to characterize a given application by classifying the sensitive voice commands it accepts. We develop a natural language processing tool that classifies a given voice command from two dimensions: (1) whether the voice command is designed to insert action or retrieve information; (2) whether the command is sensitive or nonsensitive. The tool combines a deep neural network and a keyword-based model, and uses Active Learning to reduce the manual labeling effort. The sensitivity classification is based on a user study (N=404) where we measure the perceived sensitivity of voice commands. A ground-truth evaluation shows that our tool achieves over 95% of accuracy for both types of classifications. We apply this tool to analyze 77,957 Amazon Alexa applications and 4,813 Google Home applications (198,199 voice commands from Amazon Alexa, 13,644 voice commands from Google Home) over two years (2018-2019). In total, we identify 19,263 sensitive “action injection” commands and 5,352 sensitive “information retrieval” commands. These commands are from 4,596 applications (5.55% out of all applications), most of which belong to the “smart home” category. While the percentage of sensitive applications is small, we show the percentage is increasing over time from 2018 to 2019.