The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance

The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance
复制标题

部分欺骗数据库及短假语音段检测对策

DOI:
10.1109/taslp.2022.3233236
复制
发表时间:
2022-04
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Lin Zhang;Xin Wang;Erica Cooper;Nicholas Evans;J. Yamagishi
Lin Zhang;Xin Wang;Erica Cooper;Nicholas Evans;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Lin Zhang;Xin Wang;Erica Cooper;Nicholas Evans;J. Yamagishi

文献摘要

相似文献

自动说话人确认容易受到各种操纵和欺骗,如文本到语音合成,语音转换,重放,篡改,对抗性攻击,等等。我们考虑一种新的欺骗场景称为“部分欺骗”(PS),合成或转换的语音片段嵌入到一个真正的话语。虽然现有的对策(CM)可以检测完全欺骗的话语,有必要为他们的适应或扩展到PS的情况。我们提出了各种改进,以构建一个更准确的CM,可以检测和定位短生成的欺骗语音段在更精细的时间分辨率。首先,我们引入新开发的自监督预训练模型作为增强的特征提取器。其次,我们扩展我们的PartialSpoof数据库添加段标签的各种时间分辨率。由于短欺骗的语音段被嵌入攻击者是可变长度的,六个不同的时间分辨率被认为是,从短至20毫秒到大至640毫秒。第三,我们提出了一个新的CM,使在不同的时间分辨率,以及话语级标签的段级标签的同时使用执行话语和段级检测在同一时间。我们还表明,建议的CM是能够检测欺骗在话语水平与低错误率在PS的情况下,以及在相关的逻辑访问(LA)的情况下。PartialSpoof数据库和ASVspoof 2019 LA数据库上的话语级检测的等错误率分别为0.77%和0.90%。
Automatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and so on. We consider a new spoofing scenario called “Partial Spoof” (PS) in which synthesized or transformed speech segments are embedded into a bona fide utterance. While existing countermeasures (CMs) can detect fully spoofed utterances, there is a need for their adaptation or extension to the PS scenario. We propose various improvements to construct a significantly more accurate CM that can detect and locate short-generated spoofed speech segments at finer temporal resolutions. First, we introduce newly developed self-supervised pre-trained models as enhanced feature extractors. Second, we extend our PartialSpoof database by adding segment labels for various temporal resolutions. Since the short spoofed speech segments to be embedded by attackers are of variable length, six different temporal resolutions are considered, ranging from as short as 20 ms to as large as 640 ms. Third, we propose a new CM that enables the simultaneous use of the segment-level labels at different temporal resolutions as well as utterance-level labels to execute utterance- and segment-level detection at the same time. We also show that the proposed CM is capable of detecting spoofing at the utterance level with low error rates in the PS scenario as well as in a related logical access (LA) scenario. The equal error rates of utterance-level detection on the PartialSpoof database and ASVspoof 2019 LA database were 0.77 and 0.90%, respectively.