Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data

Transforming Spectrum and Prosody for Emotional Voice Conversion with Non-Parallel Training Data
复制标题

DOI:
10.21437/odyssey.2020-33
复制
发表时间:
2020-02
期刊:
--
影响因子:
--
通讯作者:
Kun Zhou;Berrak Sisman;Haizhou Li
Kun Zhou;Berrak Sisman;Haizhou Li
中科院分区:
其他
文献类型:
--
作者:
Kun Zhou;Berrak Sisman;Haizhou Li

文献摘要

被引文献

相似文献

情感语音转换的目的是在保留说话人身份和语言内容的前提下,通过对频谱和韵律的转换来改变言语的情感模式。许多研究需要不同情绪模式之间的平行语言数据,这在现实生活中是不现实的。此外,他们经常用一个简单的线性变换来模拟基频(F0)的转换。由于F0是语调的一个关键方面,具有层次性,我们认为使用小波变换在不同的时间尺度上对F0进行建模更合适。我们提出了一个CycleGAN网络,通过使用对抗和循环一致性损失同时学习正反映射,从非并行训练数据中找到最优伪对。我们还研究了使用连续小波变换(CWT)将F0分解为描述不同时间分辨率下语音韵律的十个时间尺度,以实现有效的F0转换。实验结果表明,我们提出的框架在客观和主观评价方面都优于基线。
Emotional voice conversion aims to convert the spectrum and prosody to change the emotional patterns of speech, while preserving the speaker identity and linguistic content. Many studies require parallel speech data between different emotional patterns, which is not practical in real life. Moreover, they often model the conversion of fundamental frequency (F0) with a simple linear transform. As F0 is a key aspect of intonation that is hierarchical in nature, we believe that it is more adequate to model F0 in different temporal scales by using wavelet transform. We propose a CycleGAN network to find an optimal pseudo pair from non-parallel training data by learning forward and inverse mappings simultaneously using adversarial and cycle-consistency losses. We also study the use of continuous wavelet transform (CWT) to decompose F0 into ten temporal scales, that describes speech prosody at different time resolution, for effective F0 conversion. Experimental results show that our proposed framework outperforms the baselines both in objective and subjective evaluations.