Privacy against Real-Time Speech Emotion Detection via Acoustic Adversarial Evasion of Machine Learning

Privacy against Real-Time Speech Emotion Detection via Acoustic Adversarial Evasion of Machine Learning
复制标题

DOI:
10.1145/3610887
复制
发表时间:
2022-11
影响因子:
--
通讯作者:
Brian Testa;Yi Xiao;Harshit Sharma;Avery Gump;Asif Salekin
Brian Testa;Yi Xiao;Harshit Sharma;Avery Gump;Asif Salekin
中科院分区:
--
文献类型:
--
作者:
Brian Testa;Yi Xiao;Harshit Sharma;Avery Gump;Asif Salekin

文献摘要

相似文献

Amazon Echo和Google Home等智能扬声器语音助理(VA)由于与智能家居设备和物联网(IoT)技术的无缝集成而被广泛采用。这些VA服务引起了隐私问题,特别是由于他们可以访问我们的演讲。这项工作考虑了这样一个用例:通过语音情感识别(SER)对用户的情感进行不负责任和未经授权的监视。本文介绍了DARE-GP,这是一种解决方案,它可以创建加性噪声来掩盖用户的情感信息,同时保留语音中与转录相关的部分。DARE-GP通过使用约束遗传编程方法来学习描述目标用户情感内容的频谱频率特征,然后生成提供这种隐私保护的通用对抗性音频扰动。与现有的作品不同,DARE-GP提供:a)实时保护以前听不到的话语,B)对以前看不见的黑盒SER分类器,c)同时保护语音转录,和d)在现实的声学环境中这样做。此外,这种规避对于知识渊博的对手所采用的防御是鲁棒的。这项工作中的评估最终针对两个现成的商业智能扬声器进行声学评估,这些扬声器使用与唤醒词系统集成的小尺寸(raspberry pi)来评估其真实世界实时部署的功效。
Smart speaker voice assistants (VAs) such as Amazon Echo and Google Home have been widely adopted due to their seamless integration with smart home devices and the Internet of Things (IoT) technologies. These VA services raise privacy concerns, especially due to their access to our speech. This work considers one such use case: the unaccountable and unauthorized surveillance of a user's emotion via speech emotion recognition (SER). This paper presents DARE-GP, a solution that creates additive noise to mask users' emotional information while preserving the transcription-relevant portions of their speech. DARE-GP does this by using a constrained genetic programming approach to learn the spectral frequency traits that depict target users' emotional content, and then generating a universal adversarial audio perturbation that provides this privacy protection. Unlike existing works, DARE-GP provides: a) real-time protection of previously unheard utterances, b) against previously unseen black-box SER classifiers, c) while protecting speech transcription, and d) does so in a realistic, acoustic environment. Further, this evasion is robust against defenses employed by a knowledgeable adversary. The evaluations in this work culminate with acoustic evaluations against two off-the-shelf commercial smart speakers using a small-form-factor (raspberry pi) integrated with a wake-word system to evaluate the efficacy of its real-world, real-time deployment.