Emotional voice conversion using deep neural networks with MCC and F0 features

Emotional voice conversion using deep neural networks with MCC and F0 features
复制标题

DOI:
10.1109/icis.2016.7550889
复制
发表时间:
2016-06
期刊:
2016 IEEE/ACIS 15th International Conference on Computer and Information Science (ICIS)
影响因子:
--
通讯作者:
Zhaojie Luo;T. Takiguchi;Y. Ariki
Zhaojie Luo;T. Takiguchi;Y. Ariki
中科院分区:
其他
文献类型:
--
作者:
Zhaojie Luo;T. Takiguchi;Y. Ariki

文献摘要

被引文献

相似文献

人工神经网络是语音转换任务中训练特征的最重要的模型之一。通常,神经网络(NN)在处理低维F0特征时效果不佳,这导致基于神经网络的Mel倒谱系数(MCC)训练方法性能不突出。然而,F0可以鲁棒地表示各种韵律信号(例如,情感韵律)。在这项研究中,我们提出了一种有效的方法,基于神经网络来训练归一化段F0特征(NSF0)的情感韵律转换。同时,该方法采用深度信念网络(DBNs)训练语音转换的频谱特征。通过使用这些方法,所提出的方法可以改变频谱和韵律的情绪的声音在同一时间。实验结果表明,该方法在语音情感转换方面优于其他方法。
An artificial neural network is one of the most important models for training features in a voice conversion task. Typically, Neural Networks (NNs) are not effective in processing low-dimensional F0 features, thus this causes that the performance of those methods based on neural networks for training Mel Cepstral Coefficients (MCC) are not outstanding. However, F0 can robustly represent various prosody signals (e.g., emotional prosody). In this study, we propose an effective method based on the NNs to train the normalized-segment-F0 features (NSF0) for emotional prosody conversion. Meanwhile, the proposed method adopts deep belief networks (DBNs) to train spectrum features for voice conversion. By using these approaches, the proposed method can change the spectrum and the prosody for the emotional voice at the same time. Moreover, the experimental results show that the proposed method outperforms other state-of-the-art methods for voice emotional conversion.