Hello, Is It Me You're Looking For?: Differentiating Between Human and Electronic Speakers for Voice Interface Security

Hello, Is It Me You're Looking For?: Differentiating Between Human and Electronic Speakers for Voice Interface Security
复制标题

DOI:
10.1145/3212480.3212505
复制
发表时间:
2018-06
期刊:
Proceedings of the 11th ACM Conference on Security & Privacy in Wireless and Mobile Networks
影响因子:
--
通讯作者:
Logan Blue;Luis Vargas;Patrick Traynor
Logan Blue;Luis Vargas;Patrick Traynor
中科院分区:
其他
文献类型:
--
作者:
Logan Blue;Luis Vargas;Patrick Traynor

文献摘要

相似文献

语音接口越来越多地集成到各种物联网(IoT)设备中。这样的系统可以极大地简化用户与具有有限显示器的设备之间的交互。不幸的是,语音接口也创造了新的利用机会。具体地,在实现语音接口的系统的范围内的任何发声设备(例如,智能电视、互联网连接的设备等)可能潜在地导致这些系统执行违背其所有者的期望的操作(例如,解锁门、进行未经授权的购买等)。我们通过开发一种技术来识别人类和电子扬声器创建的音频的根本差异来解决这个问题。我们将次低音过激励,或在人类声音范围之外但现代扬声器设计所固有的显著低频信号的存在,确定为这两个源之间的强区分因素。在识别出这种现象后,我们展示了它在安静环境中以100%/1.72%TPR/FPR防止对抗性请求、重放音频和隐藏命令的用途。在这样做的过程中,我们证明,通过附近的音频设备注入的命令可以有效地删除语音接口。
Voice interfaces are increasingly becoming integrated into a variety of Internet of Things (IoT) devices. Such systems can dramatically simplify interactions between users and devices with limited displays. Unfortunately voice interfaces also create new opportunities for exploitation. Specifically any sound-emitting device within range of the system implementing the voice interface (e.g., a smart television, an Internet-connected appliance, etc) can potentially cause these systems to perform operations against the desires of their owners (e.g., unlock doors, make unauthorized purchases, etc). We address this problem by developing a technique to recognize fundamental differences in audio created by humans and electronic speakers. We identify sub-bass over-excitation, or the presence of significant low frequency signals that are outside of the range of human voices but inherent to the design of modern speakers, as a strong differentiator between these two sources. After identifying this phenomenon, we demonstrate its use in preventing adversarial requests, replayed audio, and hidden commands with a 100%/1.72% TPR/FPR in quiet environments. In so doing, we demonstrate that commands injected via nearby audio devices can be effectively removed by voice interfaces.