Privacy-preserving Liveness Detection for Securing Smart Voice Interfaces

Privacy-preserving Liveness Detection for Securing Smart Voice Interfaces
复制标题

DOI:
10.1109/tdsc.2023.3319833
复制
发表时间:
2023
影响因子:
7.3
通讯作者:
Yan Meng;Jiachun Li;Haojin Zhu;Yuan Tian;Jiming Chen
Yan Meng;Jiachun Li;Haojin Zhu;Yuan Tian;Jiming Chen
中科院分区:
计算机科学2区
文献类型:
--
作者:
Yan Meng;Jiachun Li;Haojin Zhu;Yuan Tian;Jiming Chen

文献摘要

相似文献

-智能扬声器广泛用作智能系统的主要用户界面,包括智能家居和工业物联网。然而,它们很容易受到语音欺骗攻击,从而导致恶意命令执行或隐私信息泄露。被动活跃度检测通过分析收集的音频来阻止语音欺骗,而不是部署传感器来区分真人和欺骗性语音,已经引起了越来越多的关注。但现有的方案要么在环境因素变化下面临性能下降,要么需要用户保持固定的手势,这限制了它们在现实世界场景中的部署。此外,智能扬声器的空间分布特性导致为所有相关用户构建通用分类器非常繁琐,并增加了隐私泄露问题。为了应对上述挑战,我们提出了L Ive A RRAY,一个高效、轻量级、保护隐私的被动活性检测系统。L Ive A RRAY利用了一种新的活体特征-阵列指纹,该特征利用智能扬声器固有的麦克风阵列来提高活体检测的准确性。L的S进一步采用了基于联邦学习的体系结构,减少了分类器构建过程中的数据集收集开销,消除了数据传输过程中潜在的隐私泄露。实验结果表明,L Ive A RRAY算法的准确率达到了99.16%,明显优于现有的被动算法。
—Smart speakers are widely used as the primary user interface in intelligent systems, including smart homes and industrial IoT. However, they are vulnerable to voice spoofing attacks which result in malicious command execution or privacy information leakage. Passive liveness detection, which thwarts voice spoofing via analyzing the collected audio rather than deploying sensors to distinguish between live-human and spoofing voices, has drawn increasing attention. But existing schemes either face performance degradation under environmental factor changes or require the user to keep fixed gestures, which limit their deployment in real-world scenarios. Besides, the space distributed property of smart speakers causes building a universal classifier for all involved users to be cumbersome and increases privacy leakage issues. To address the challenges mentioned above, we propose L IVE A RRAY , an efficient, lightweight, and privacy-preserving passive liveness detection system. L IVE A RRAY exploits a novel liveness feature, array fingerprint, which utilizes the microphone array inherently adopted by the smart speaker to improve the accuracy of liveness detection. L IVE A RRAY ’s further employs the federated learning-based architecture to reduce the dataset collection overhead during classifier building and eliminate the potential privacy leakage during data transmission. Experimental results show that L IVE A RRAY achieves an accuracy of 99.16%, which is superior to existing passive schemes.