A Light-Weight Replay Detection Framework For Voice Controlled IoT Devices

A Light-Weight Replay Detection Framework For Voice Controlled IoT Devices
复制标题

DOI:
10.1109/jstsp.2020.2999828
复制
发表时间:
2020-08-01
影响因子:
7.5
通讯作者:
Irtaza, Aun
Irtaza, Aun
中科院分区:
工程技术1区
文献类型:
--
作者:
Malik, Khalid Mahmood;Javed, Ali;Irtaza, Aun

文献摘要

被引文献

相似文献

语音控制设备(VCD)的数量不断增加,如Google Home、Amazon Alexa等,导致了家用电器、智能设备和下一代汽车等的自动化。然而,VCD和语音激活服务(如聊天机器人)容易受到音频重放攻击。我们对VCD的漏洞分析表明,这些回放可以在多跳场景中被利用,以恶意访问物联网连接的设备/节点。为了保护这些VCD和声控服务,迫切需要开发可靠且计算高效的解决方案来检测重播攻击。本文将重放攻击建模为引入高次谐波失真的非线性过程。为了检测这些谐波失真,我们提出了能够捕获重放攻击引起的失真的声学三值模式-伽马酮倒谱系数(ATP-GTCC)特征。利用纠错输出编码模型,利用提出的ATP-GTCC特征空间训练多类支持向量机分类器,并对其进行语音重放攻击检测。在ASVspoof 2019数据集和我们自己创建的语音欺骗检测语料库(VSDC)上对所提出的框架的性能进行了评估,该语料库包括真实的一阶重放(重放一次)和二阶重放(重放两次)音频记录。实验结果表明,该音频重放检测框架能够可靠地检测出一阶和二阶重放攻击,适用于资源受限的设备。
The growing number of voice-controlled devices (VCDs), i.e. Google Home, Amazon Alexa, etc., has resulted in automation of home appliances, smart gadgets, and next generation vehicles, etc. However, VCDs and voice-activated services i.e. chatbots are vulnerable to audio replay attacks. Our vulnerability analysis of VCDs shows that these replays could be exploited in multi-hop scenarios to maliciously access the devices/nodes attached to the Internet of Things. To protect these VCDs and voice-activated services, there is an urgent need to develop reliable and computationally efficient solutions to detect the replay attacks. This paper models replay attacks as a nonlinear process that introduces higher-order harmonic distortions. To detect these harmonic distortions, we propose the acoustic ternary patterns-gammatone cepstral coefficient (ATP-GTCC) features that are capable of capturing distortions due to replay attacks. Error correcting output codes model is used to train a multi-class SVM classifier using the proposed ATP-GTCC feature space and tested for voice replay attack detection. Performance of the proposed framework is evaluated on ASVspoof 2019 dataset, and our own created voice spoofing detection corpus (VSDC) consisting of bona-fide, first-order replay (replayed once), and second-order replay (replayed twice) audio recordings. Experimental results signify that the proposed audio replay detection framework reliably detects both first and second-order replay attacks and can be used in resource constrained devices.