Unacceptable, where is my privacy? Exploring Accidental Triggers of Smart Speakers

Unacceptable, where is my privacy? Exploring Accidental Triggers of Smart Speakers
复制标题

DOI:
--
复制
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Lea Schönherr;Maximilian Golla;Thorsten Eisenhofer;Jan Wiele;D. Kolossa;Thorsten Holz
Lea Schönherr;Maximilian Golla;Thorsten Eisenhofer;Jan Wiele;D. Kolossa;Thorsten Holz
中科院分区:
其他
文献类型:
--
作者:
Lea Schönherr;Maximilian Golla;Thorsten Eisenhofer;Jan Wiele;D. Kolossa;Thorsten Holz

文献摘要

被引文献

相似文献

像亚马逊的Alexa、谷歌的Assistant或苹果的Siri这样的语音助手已经成为数百万家庭中智能扬声器的主要(语音)接口。出于隐私的原因,这些扬声器在将音频流上传到云中进行进一步处理之前,会分析环境中的每一种声音,以获得各自的唤醒单词,如“Alexa”或“嘿Siri”。之前的研究报告了不准确的唤醒单词检测,可以用类似的单词或发音来欺骗它,比如“可卡因面条”,而不是“OK Google”。在本文中,我们对这种意外触发进行了全面的分析,即不应该触发语音助手但却触发了的声音。更具体地说,我们使用电视节目、新闻和其他类型的音频数据集等日常媒体,自动化了查找意外触发因素的过程,并测量了来自8个不同制造商的11个智能扬声器的流行度。为了系统地检测意外触发因素,我们描述了一种方法,使用发音词典和基于电话的加权Lvenshtein距离来人工制作此类触发因素。总体而言,我们已经发现了数百个意外触发因素。此外,我们还探讨了潜在的性别和语言偏见,并分析了其再现性。最后,我们讨论了意外触发对隐私的影响,并探索了减少和限制其对用户隐私影响的对策。为了促进对这些误导机器学习模型的声音的进一步研究,我们发布了一个包含1000多个经过验证的触发器的数据集作为研究人工制品。
Voice assistants like Amazon's Alexa, Google's Assistant, or Apple's Siri, have become the primary (voice) interface in smart speakers that can be found in millions of households. For privacy reasons, these speakers analyze every sound in their environment for their respective wake word like ''Alexa'' or ''Hey Siri,'' before uploading the audio stream to the cloud for further processing. Previous work reported on the inaccurate wake word detection, which can be tricked using similar words or sounds like ''cocaine noodles'' instead of ''OK Google.'' In this paper, we perform a comprehensive analysis of such accidental triggers, i.,e., sounds that should not have triggered the voice assistant, but did. More specifically, we automate the process of finding accidental triggers and measure their prevalence across 11 smart speakers from 8 different manufacturers using everyday media such as TV shows, news, and other kinds of audio datasets. To systematically detect accidental triggers, we describe a method to artificially craft such triggers using a pronouncing dictionary and a weighted, phone-based Levenshtein distance. In total, we have found hundreds of accidental triggers. Moreover, we explore potential gender and language biases and analyze the reproducibility. Finally, we discuss the resulting privacy implications of accidental triggers and explore countermeasures to reduce and limit their impact on users' privacy. To foster additional research on these sounds that mislead machine learning models, we publish a dataset of more than 1000 verified triggers as a research artifact.