Towards More Robust Keyword Spotting for Voice Assistants

Towards More Robust Keyword Spotting for Voice Assistants
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
Proceedings of the 36th Annual Computer Security Applications Conference
影响因子:
--
通讯作者:
Shimaa Ahmed;Ilia Shumailov;Nicolas Papernot;Kassem Fawaz
Shimaa Ahmed;Ilia Shumailov;Nicolas Papernot;Kassem Fawaz
中科院分区:
其他
文献类型:
--
作者:
Shimaa Ahmed;Ilia Shumailov;Nicolas Papernot;Kassem Fawaz

文献摘要

相似文献

语音助手依靠关键字识别(KWS)来处理人类发出的语音命令:命令前面有一个关键字,如“Alexa”或“OK Google”,必须找到这些关键字才能激活语音助手。通常,关键字识别有两个方面:设备上的模型首先识别关键字,然后产生的语音样本触发第二个云上模型,后者验证和处理激活。在这项工作中,我们探讨了这在两种威胁模型下引发的重大隐私和安全问题。首先,我们的实验表明,意外激活会导致长达一分钟的语音记录被上传到云中。其次,我们通过敌意的例子验证了攻击者可以系统地触发误激活,这暴露了与语音助手连接的服务的完整性和可用性。我们提出了EKOS(Ensymble For Keyword Spotting),它利用KWS任务的语义来防御意外和恶意激活。EKOS在训练和推理时结合了来自声学环境的空间冗余,以最大限度地减少导致意外激活的分布漂移。它还利用语音的物理特性--它在不同谐波下的冗余--部署针对不同谐波训练的模型集合,并可证明地迫使对手修改更多的频谱以获得对手示例。我们的评估表明,EKOS增加了对抗性激活的成本,同时保持了自然的准确性。我们通过在商用设备和商用语音助手上的空中实验验证了EKOS的性能;我们发现EKOS在非对抗性环境下提高了KWS任务的精度。
Voice assistants rely on keyword spotting (KWS) to process vocal commands issued by humans: commands are prepended with a keyword, such as “Alexa” or “Ok Google,” which must be spotted to activate the voice assistant. Typically, keyword spotting is two-fold: an on-device model first identifies the keyword, then the resulting voice sample triggers a second on-cloud model which verifies and processes the activation. In this work, we explore the significant privacy and security concerns that this raises under two threat models. First, our experiments demonstrate that accidental activations result in up to a minute of speech recording being uploaded to the cloud. Second, we verify that adversaries can systematically trigger misactivations through adversarial examples, which exposes the integrity and availability of services connected to the voice assistant. We propose EKOS (Ensemble for KeywOrd Spotting) which leverages the semantics of the KWS task to defend against both accidental and adversarial activations. EKOS incorporates spatial redundancy from the acoustic environment at training and inference time to minimize distribution drifts responsible for accidental activations. It also exploits a physical property of speech—its redundancy at different harmonics—to deploy an ensemble of models trained on different harmonics and provably force the adversary to modify more of the frequency spectrum to obtain adversarial examples. Our evaluation shows that EKOS increases the cost of adversarial activations, while preserving the natural accuracy. We validate the performance of EKOS with over-the-air experiments on commodity devices and commercial voice assistants; we find that EKOS improves the precision of the KWS task in non-adversarial settings.