Voice spoofing detection corpus for single and multi-order audio replays

Voice spoofing detection corpus for single and multi-order audio replays
复制标题

DOI:
10.1016/j.csl.2020.101132
复制
发表时间:
2021-01-01
影响因子:
4.3
通讯作者:
Malik, Hafiz
Malik, Hafiz
中科院分区:
计算机科学3区
文献类型:
--
作者:
Baumann, Roland;Malik, Khalid Mahmood;Malik, Hafiz

文献摘要

被引文献

相似文献

现代语音控制设备(vcd)的发展彻底改变了物联网(IoT),并通过语音命令增加了智能家居、个性化和家庭自动化的实现。这些vcd可以在物联网驱动的环境中被利用来产生各种欺骗攻击,包括重放攻击链(即多阶重放攻击)。现有数据集,如ASVspoof 2017、ASVspoof 2019和ReMASC仅包含一阶重播记录(即重播一次);因此,它们不能提供能够检测多阶重放攻击的反欺骗算法的评估。此外,像ASVspoof 2017和ASVspoof 2019这样的大规模数据集没有捕捉到麦克风阵列的特征,而麦克风阵列是现代vcd的基本特征。因此,需要一种不同的重播欺骗检测语料库,该语料库由针对真实语音样本的多阶重播记录组成。本文提出了一种新的语音欺骗检测语料库(VSDC)来评估多阶重放抗欺骗方法的性能。所提出的VSDC由一阶(即重播一次)和二阶重播(即重播两次)针对真实音频记录的样本组成。我们确保在环境、录音和回放设备、扬声器、配置、回放场景等方面创建多样化的重放欺骗检测语料库。更具体地说,我们使用了35个麦克风,25种不同的录音配置,60种不同的播放配置,用于一阶和二阶重播,共生成属于19个扬声器的14,050个样本。此外,本文提出的VSDC也可用于评估说话人验证系统在独立说话人验证方面的性能。据我们所知,这是第一个公开可用的重播欺骗检测语料库,由一阶和二阶重播样本组成。实验结果表明,本文提出的VSDC在多阶重放攻击和不同条件下的抗欺骗性能评估方面是有效的。(C) 2020 Elsevier Ltd.版权所有。
The evolution of modern voice-controlled devices (VCDs) has revolutionized the Internet of Things (IoT) and resulted in the increased realization of smart homes, personalization, and home automation through voice commands. These VCDs can be exploited in IoT driven environments to generate various spoofing attacks, including the chaining of replay attacks (i.e. multi-order replay attacks). Existing datasets like ASVspoof 2017, ASVspoof 2019, and ReMASC contain only first-order replay recordings (i.e. replayed once); therefore, they cannot offer evaluation of anti-spoofing algorithms capable of detecting multi-order replay attacks. Additionally, large-scale datasets like ASVspoof 2017 and ASVspoof 2019 do not capture the characteristics of microphone arrays, which are an essential characteristic of modern VCDs. Therefore, there exists a need for a diverse replay spoofing detection corpus that consists of multi-order replay recordings against bona fide voice samples. This paper presents a novel voice spoofing detection corpus (VSDC) to evaluate the performance of multi-order replay anti-spoofing methods. The proposed VSDC consists of first-order (i.e. replayed once) and second-order replay (i.e. replayed twice) samples against the bona fide audio recordings. We ensured to create a diverse replay spoofing detection corpus in terms of environments, recording and playback devices, speakers, configurations, replay scenarios, etc. More specifically, we used 35 microphones, 25 different recording configurations, and 60 different playback configurations for first- and second-order replays to generate a total of 14,050 samples belonging to 19 speakers. Additionally, the proposed VSDC can also be used to evaluate the performance of speaker verification systems in terms of independent speaker verification. To the best of our knowledge, this is the first publicly available replay spoofing detection corpus comprised of first and second-order replay samples. Experimental results signify the effectiveness of the proposed VSDC in terms of evaluating the performance of anti-spoofing methods under multi-order replay attacks and diverse conditions. (C) 2020 Elsevier Ltd. All rights reserved.