Voice Conversion Challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion

Voice Conversion Challenge 2020: Intra-lingual semi-parallel and cross-lingual voice conversion
复制标题

DOI:
10.21437/vcc_bc.2020-14
复制
发表时间:
2020-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Yi Zhao;Wen-Chin Huang;Xiaohai Tian;J. Yamagishi;Rohan Kumar Das;T. Kinnunen;Zhenhua Ling;T. Toda
Yi Zhao;Wen-Chin Huang;Xiaohai Tian;J. Yamagishi;Rohan Kumar Das;T. Kinnunen;Zhenhua Ling;T. Toda
中科院分区:
其他
文献类型:
--
作者:
Yi Zhao;Wen-Chin Huang;Xiaohai Tian;J. Yamagishi;Rohan Kumar Das;T. Kinnunen;Zhenhua Ling;T. Toda

文献摘要

被引文献

相似文献

语音转换挑战是一项两年一次的科学活动,旨在比较和理解基于公共数据集的不同语音转换(VC)系统。2020年,我们组织了第三届挑战赛,并构建和分发了一个新的数据库,用于两项任务:语内半并行和跨语言VC。经过两个月的挑战期,我们收到了33份意见书,其中包括建立在数据库上的3条基线。从众包听力测试的结果中,我们观察到,由于先进的深度学习方法,VC方法发展迅速。特别是,在语内半平行VC任务中,几个系统的说话人相似度得分与目标说话人一样高。然而,我们确认,对于同样的任务,它们都没有达到人类的自然水平。正如预期的那样,跨语言转换任务是一个更困难的任务,总体自然度和相似度得分低于语内转换任务。然而,我们观察到令人鼓舞的结果,最佳系统的MOS评分高于4.0。我们还展示了一些额外的分析结果,以帮助更好地理解跨语言VC。
The voice conversion challenge is a bi-annual scientific event held to compare and understand different voice conversion (VC) systems built on a common dataset. In 2020, we organized the third edition of the challenge and constructed and distributed a new database for two tasks, intra-lingual semi-parallel and cross-lingual VC. After a two-month challenge period, we received 33 submissions, including 3 baselines built on the database. From the results of crowd-sourced listening tests, we observed that VC methods have progressed rapidly thanks to advanced deep learning methods. In particular, speaker similarity scores of several systems turned out to be as high as target speakers in the intra-lingual semi-parallel VC task. However, we confirmed that none of them have achieved human-level naturalness yet for the same task. The cross-lingual conversion task is, as expected, a more difficult task, and the overall naturalness and similarity scores were lower than those for the intra-lingual conversion task. However, we observed encouraging results, and the MOS scores of the best systems were higher than 4.0. We also show a few additional analysis results to aid in understanding cross-lingual VC better.