Voice Presentation Attack Detection through Text-Converted Voice Command Analysis

Voice Presentation Attack Detection through Text-Converted Voice Command Analysis
复制标题

DOI:
10.1145/3290605.3300828
复制
发表时间:
2019-05
期刊:
Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems
影响因子:
--
通讯作者:
Il-Youp Kwak;J. Huh;S. Han;Iljoo Kim;J. Yoon
Il-Youp Kwak;J. Huh;S. Han;Iljoo Kim;J. Yoon
中科院分区:
其他
文献类型:
--
作者:
Il-Youp Kwak;J. Huh;S. Han;Iljoo Kim;J. Yoon

文献摘要

被引文献

相似文献

语音助手正在迅速升级,以支持高级的安全关键命令,例如解锁设备,检查电子邮件和进行支付。在本文中,我们探讨了使用用户的文本转换的语音命令话语作为分类特征,以帮助识别用户的真正的命令,并检测可疑的命令的可行性。为了保持高检测精度,我们的方法从全局训练的攻击检测模型(立即可用于新用户)开始,并逐渐切换到针对目标用户的话语模式定制的用户特定模型。为了评估准确性,我们使用了一个真实的语音助手数据集,其中包括从260万用户那里收集的大约3460万条语音命令。我们的评估结果表明,当使用最佳阈值时,该方法能够实现约3.4%的等错误率(EER),检测到95.7%的攻击。对于那些经常使用安全关键(类似攻击)命令的人,我们仍然可以实现低于5%的EER。
Voice assistants are quickly being upgraded to support advanced, security-critical commands such as unlocking devices, checking emails, and making payments. In this paper, we explore the feasibility of using users' text-converted voice command utterances as classification features to help identify users' genuine commands, and detect suspicious commands. To maintain high detection accuracy, our approach starts with a globally trained attack detection model (immediately available for new users), and gradually switches to a user-specific model tailored to the utterance patterns of a target user. To evaluate accuracy, we used a real-world voice assistant dataset consisting of about 34.6 million voice commands collected from 2.6 million users. Our evaluation results show that this approach is capable of achieving about 3.4% equal error rate (EER), detecting 95.7% of attacks when an optimal threshold value is used. As for those who frequently use security-critical (attack-like) commands, we still achieve EER below 5%.