Emotional voice conversion using deep neural networks with MCC and F0 features
Emotional voice conversion using deep neural networks with MCC and F0 features
复制标题
DOI:
10.1109/icis.2016.7550889
复制
发表时间:
2016-06
期刊:
影响因子:
--
通讯作者:
Zhaojie Luo;T. Takiguchi;Y. Ariki
中科院分区:
文献类型:
--
作者:
Zhaojie Luo;T. Takiguchi;Y. Ariki
An artificial neural network is one of the most important models for training features in a voice conversion task. Typically, Neural Networks (NNs) are not effective in processing low-dimensional F0 features, thus this causes that the performance of those methods based on neural networks for training Mel Cepstral Coefficients (MCC) are not outstanding. However, F0 can robustly represent various prosody signals (e.g., emotional prosody). In this study, we propose an effective method based on the NNs to train the normalized-segment-F0 features (NSF0) for emotional prosody conversion. Meanwhile, the proposed method adopts deep belief networks (DBNs) to train spectrum features for voice conversion. By using these approaches, the proposed method can change the spectrum and the prosody for the emotional voice at the same time. Moreover, the experimental results show that the proposed method outperforms other state-of-the-art methods for voice emotional conversion.