A Spoofing Benchmark for the 2018 Voice Conversion Challenge: Leveraging from Spoofing Countermeasures for Speech Artifact Assessment

A Spoofing Benchmark for the 2018 Voice Conversion Challenge: Leveraging from Spoofing Countermeasures for Speech Artifact Assessment
复制标题

DOI:
10.21437/odyssey.2018-27
复制
发表时间:
2018-04
期刊:
--
影响因子:
--
通讯作者:
T. Kinnunen;Jaime Lorenzo-Trueba;J. Yamagishi;T. Toda;D. Saito;F. Villavicencio;Zhenhua Ling
T. Kinnunen;Jaime Lorenzo-Trueba;J. Yamagishi;T. Toda;D. Saito;F. Villavicencio;Zhenhua Ling
中科院分区:
其他
文献类型:
--
作者:
T. Kinnunen;Jaime Lorenzo-Trueba;J. Yamagishi;T. Toda;D. Saito;F. Villavicencio;Zhenhua Ling

文献摘要

相似文献

语音转换的目的是在不改变说话人语音内容的前提下,实现说话人特征的转换。由于训练数据的限制和建模的不完善,很难在不引入处理伪像的情况下实现可信的说话人模仿;因此,VC的性能评估通常涉及说话人相似性和人类小组的质量评估。作为一个耗时、昂贵且不可复制的过程,它阻碍了新VC技术的快速原型制作。我们使用另一种客观的方法来解决伪影评估问题,该方法利用了自动说话人验证的欺骗对策(CM)的先前工作。其中,CM用于拒绝“假”输入,例如重放的、合成的或转换的语音,但它们用于自动语音伪影评估的潜力仍然未知。本研究旨在填补这一空白。作为对2018年语音转换挑战赛(VCC'18)数据主观结果的补充,我们配置了一个标准的恒定Q倒谱系数CM来量化处理伪影的程度。CM的等错误率(EER),VC样本与真实的人类语音的混淆性指标,作为我们的伪影测量。识别了VCC'18条目的两个集群:具有可检测伪影的低质量条目(低EER)和具有较少伪影的较高质量条目。然而,VCC'18系统中没有一个是完美的:所有EER都<30%(“理想”值为50%)。我们的初步研究结果表明,CM在其原始应用程序之外的潜力,作为一种补充优化和基准测试工具,以提高VC技术。
Voice conversion (VC) aims at conversion of speaker characteristic without altering content. Due to training data limitations and modeling imperfections, it is difficult to achieve believable speaker mimicry without introducing processing artifacts; performance assessment of VC, therefore, usually involves both speaker similarity and quality evaluation by a human panel. As a time-consuming, expensive, and non-reproducible process, it hinders rapid prototyping of new VC technology. We address artifact assessment using an alternative, objective approach leveraging from prior work on spoofing countermeasures (CMs) for automatic speaker verification. Therein, CMs are used for rejecting `fake' inputs such as replayed, synthetic or converted speech but their potential for automatic speech artifact assessment remains unknown. This study serves to fill that gap. As a supplement to subjective results for the 2018 Voice Conversion Challenge (VCC'18) data, we configure a standard constant-Q cepstral coefficient CM to quantify the extent of processing artifacts. Equal error rate (EER) of the CM, a confusability index of VC samples with real human speech, serves as our artifact measure. Two clusters of VCC'18 entries are identified: low-quality ones with detectable artifacts (low EERs), and higher quality ones with less artifacts. None of the VCC'18 systems, however, is perfect: all EERs are < 30 % (the `ideal' value would be 50 %). Our preliminary findings suggest potential of CMs outside of their original application, as a supplemental optimization and benchmarking tool to enhance VC technology.