Vaw-Gan For Disentanglement And Recomposition Of Emotional Elements In Speech
Vaw-Gan For Disentanglement And Recomposition Of Emotional Elements In Speech
复制标题
DOI:
10.1109/slt48900.2021.9383526
复制
发表时间:
2020-11
期刊:
影响因子:
--
通讯作者:
Kun Zhou;Berrak Sisman;Haizhou Li
中科院分区:
文献类型:
--
作者:
Kun Zhou;Berrak Sisman;Haizhou Li
Emotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition of emotional elements in speech through variational autoencoding Wasserstein generative adversarial network (VAW-GAN). We propose a speaker-dependent EVC framework based on VAW-GAN, that includes two VAW-GAN pipelines, one for spectrum conversion, and another for prosody conversion. We train a spectral encoder that disentangles emotion and prosody (F0) information from spectral features; we also train a prosodic encoder that disentangles emotion modulation of prosody (affective prosody) from linguistic prosody. At run-time, the decoder of spectral VAW-GAN is conditioned on the output of prosodic VAW-GAN. The vocoder takes the converted spectral and prosodic features to generate the target emotional speech. Experiments validate the effectiveness of our proposed method in both objective and subjective evaluations.