WearID: Low-Effort Wearable-Assisted Authentication of Voice Commands via Cross-Domain Comparison without Training

WearID: Low-Effort Wearable-Assisted Authentication of Voice Commands via Cross-Domain Comparison without Training
复制标题

DOI:
10.1145/3427228.3427259
复制
发表时间:
2020-12
期刊:
Proceedings of the 36th Annual Computer Security Applications Conference
影响因子:
--
通讯作者:
Cong Shi;Yan Wang;Yingying Chen;Nitesh Saxena;Chen Wang
Cong Shi;Yan Wang;Yingying Chen;Nitesh Saxena;Chen Wang
中科院分区:
其他
文献类型:
--
作者:
Cong Shi;Yan Wang;Yingying Chen;Nitesh Saxena;Chen Wang

文献摘要

被引文献

相似文献

由于语音输入的开放性,语音助理(VA)系统(例如,Google Home和Amazon Alexa)容易受到各种安全和隐私泄露(例如,信用卡号码、密码)的攻击,特别是当发出涉及大量购买、关键呼叫等的关键用户命令时。尽管现有的VA系统可以使用语音特征来识别用户,但它们仍然容易受到各种基于声音的攻击(例如,模仿、重放和隐藏命令攻击)。在这项工作中,我们提出了一种无需训练的语音认证系统WearID,利用音频域和振动域之间的跨域语音相似性来为不断增长的VA系统部署提供增强的安全性。特别是,当用户发出关键命令时,WearID利用用户可穿戴设备上的运动传感器来捕获振动域中的空中语音,并通过VA设备的麦克风与音频域中捕获的语音进行验证。与现有的方法相比,我们的解决方案是低工作量和隐私保护的,因为它既不需要用户的主动输入(例如,回复消息/呼叫),也不需要存储用户隐私敏感的语音样本用于训练。此外,我们的解决方案利用独特的振动感应界面及其对声音的短感应范围(例如,25厘米)来验证语音命令。检查这两个领域数据的相似性并不是一件容易的事。音频域和振动域之间的巨大采样率差距(例如,8000赫兹与200赫兹)使得很难直接比较这两个域的数据,甚至微小的数据噪声也可能被放大并导致身份验证失败。为了应对这些挑战,我们研究了两个感应域之间的复杂关系,并开发了一种基于频谱图的算法,将麦克风数据转换为频率较低的运动传感器数据,以便于跨域比较。我们进一步开发了一种用户认证方案,基于接收到的语音命令的跨域语音相似度来验证接收到的语音命令来自合法用户。我们报告了在各种可听和不可听攻击下评估WearID的广泛实验。结果表明,WearID能够在正常情况下以99.8%的准确率验证语音命令,并从包括冒充/重放攻击和隐藏语音/超声攻击在内的各种攻击中检测出97.2%的虚假语音命令。
Due to the open nature of voice input, voice assistant (VA) systems (e.g., Google Home and Amazon Alexa) are vulnerable to various security and privacy leakages (e.g., credit card numbers, passwords), especially when issuing critical user commands involving large purchases, critical calls, etc. Though the existing VA systems may employ voice features to identify users, they are still vulnerable to various acoustic-based attacks (e.g., impersonation, replay, and hidden command attacks). In this work, we propose a training-free voice authentication system, WearID, leveraging the cross-domain speech similarity between the audio domain and the vibration domain to provide enhanced security to the ever-growing deployment of VA systems. In particular, when a user gives a critical command, WearID exploits motion sensors on the user’s wearable device to capture the aerial speech in the vibration domain and verify it with the speech captured in the audio domain via the VA device’s microphone. Compared to existing approaches, our solution is low-effort and privacy-preserving, as it neither requires users’ active inputs (e.g., replying messages/calls) nor to store users’ privacy-sensitive voice samples for training. In addition, our solution exploits the distinct vibration sensing interface and its short sensing range to sound (e.g., 25cm) to verify voice commands. Examining the similarity of the two domains’ data is not trivial. The huge sampling rate gap (e.g., 8000Hz vs. 200Hz) between the audio and vibration domains makes it hard to compare the two domains’ data directly, and even tiny data noises could be magnified and cause authentication failures. To address the challenges, we investigate the complex relationship between the two sensing domains and develop a spectrogram-based algorithm to convert the microphone data into the lower-frequency “ motion sensor data” to facilitate cross-domain comparisons. We further develop a user authentication scheme to verify that the received voice command originates from the legitimate user based on the cross-domain speech similarity of the received voice commands. We report on extensive experiments to evaluate the WearID under various audible and inaudible attacks. The results show WearID can verify voice commands with 99.8% accuracy in the normal situation and detect 97.2% fake voice commands from various attacks, including impersonation/replay attacks and hidden voice/ultrasound attacks.