Singing Voice Synthesis Based on Deep Neural Networks

Singing Voice Synthesis Based on Deep Neural Networks
复制标题

DOI:
10.21437/interspeech.2016-1027
复制
发表时间:
2016-09
期刊:
--
影响因子:
--
通讯作者:
Masanari Nishimura;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda
Masanari Nishimura;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda
中科院分区:
其他
文献类型:
--
作者:
Masanari Nishimura;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda

文献摘要

相似文献

基于隐马尔可夫模型(HMM)的歌唱声音合成技术已经被提出。在这些方法中,频谱,激励,和持续时间的歌声同时建模与上下文相关的障碍和波形产生的障碍本身。然而,合成的歌声质量仍然没有达到自然歌声的质量。深度神经网络(DNN)在包括语音识别、图像识别、语音合成等在内的各个研究领域都大大改进了传统方法。基于DNN的文本到语音(TTS)合成可以合成高质量的语音。在基于DNN的TTS系统中,DNN被训练来表示从上下文特征到声学特征的映射函数,其在基于HMM的TTS系统中由决策树聚类的上下文相关的Hynth建模。在本文中,我们提出了基于DNN的歌唱声音合成,并评估其有效性。乐谱与其声学特征之间的关系由DNN在帧中建模。针对基音背景在数据库中的稀疏性,采用音符级基音归一化和线性插值技术来提取激励特征。主观实验结果表明,基于DNN的系统优于基于HMM的系统的自然度。
Singing voice synthesis techniques have been proposed based on a hidden Markov model (HMM). In these approaches, the spectrum, excitation, and duration of singing voices are simultaneously modeled with context-dependent HMMs and waveforms are generated from the HMMs themselves. However, the quality of the synthesized singing voices still has not reached that of natural singing voices. Deep neural networks (DNNs) have largely improved on conventional approaches in various research areas including speech recognition, image recognition, speech synthesis, etc. The DNN-based text-to-speech (TTS) synthesis can synthesize high quality speech. In the DNN-based TTS system, a DNN is trained to represent the mapping function from contextual features to acoustic features, which are modeled by decision tree-clustered context dependent HMMs in the HMM-based TTS system. In this paper, we propose singing voice synthesis based on a DNN and evaluate its effectiveness. The relationship between the musical score and its acoustic features is modeled in frames by a DNN. For the sparseness of pitch context in a database, a musical-note-level pitch normalization and linear-interpolation techniques are used to prepare the excitation features. Subjective experimental results show that the DNN-based system outperformed the HMM-based system in terms of naturalness.