Sequential Generation of Singing F0 Contours from Musical Note Sequences Based on WaveNet

Sequential Generation of Singing F0 Contours from Musical Note Sequences Based on WaveNet
复制标题

DOI:
10.23919/apsipa.2018.8659502
复制
发表时间:
2018-11
期刊:
2018 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC)
影响因子:
--
通讯作者:
Yusuke Wada;Ryo Nishikimi;Eita Nakamura;Katsutoshi Itoyama;Kazuyoshi Yoshii
Yusuke Wada;Ryo Nishikimi;Eita Nakamura;Katsutoshi Itoyama;Kazuyoshi Yoshii
中科院分区:
其他
文献类型:
--
作者:
Yusuke Wada;Ryo Nishikimi;Eita Nakamura;Katsutoshi Itoyama;Kazuyoshi Yoshii

文献摘要

被引文献

相似文献

本文描述了一种方法,该方法可以通过使用称为WaveNet的深度神经自回归模型从单声道音符序列(乐谱)生成连续的歌声F0轮廓。真实的F0曲线包括由诸如颤音和滑音的歌唱表达引起的复杂的时间和频率波动。虽然显式模型,如隐马尔可夫模型(HMM)经常用于表示F0动态,它是很难产生逼真的F0轮廓,由于这种模型的表示能力差。为了克服这一局限性,WaveNet被发明用于以无监督的方式对原始波形进行建模,最近被用于以监督的方式从带有歌词的乐谱生成歌唱F0轮廓。受这种尝试的启发,我们研究了WaveNet在不使用歌词信息的情况下生成唱歌F0轮廓的能力。我们的方法条件WaveNet上的音高和上下文特征的乐谱。作为一种更适合于生成F0轮廓的损失函数,我们采用了改进的交叉熵损失,该损失在对数频率轴上与目标和输出F0之间的平方误差加权。实验结果表明,这些技术提高了生成的F0轮廓的质量。
This paper describes a method that can generate a continuous F0 contour of a singing voice from a monophonic sequence of musical notes (musical score) by using a deep neural autoregressive model called WaveNet. Real F0 contours include complicated temporal and frequency fluctuations caused by singing expressions such as vibrato and portamento. Although explicit models such as hidden Markov models (HMMs) have often used for representing the F0 dynamics, it is difficult to generate realistic F0 contours due to the poor representation capability of such models. To overcome this limitation, WaveNet, which was invented for modeling raw waveforms in an unsupervised manner, was recently used for generating singing F0 contours from a musical score with lyrics in a supervised manner. Inspired by this attempt, we investigate the capability of WaveNet for generating singing F0 contours without using lyric information. Our method conditions WaveNet on pitch and contextual features of a musical score. As a loss function that is more suitable for generating F0 contours, we adopted the modified cross-entropy loss weighted with the square error between target and output F0s on the log-frequency axis. The experimental results show that these techniques improve the quality of generated F0 contours.