Transferring Neural Speech Waveform Synthesizers to Musical Instrument Sounds Generation

Transferring Neural Speech Waveform Synthesizers to Musical Instrument Sounds Generation
复制标题

DOI:
10.1109/icassp40776.2020.9053047
复制
发表时间:
2019-10
期刊:
ICASSP 2020 - 2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Yi Zhao;Xin Wang;Lauri Juvela;J. Yamagishi
Yi Zhao;Xin Wang;Lauri Juvela;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Yi Zhao;Xin Wang;Lauri Juvela;J. Yamagishi

文献摘要

相似文献

最近的神经波形合成器,如WaveNet, WaveG-low和神经源滤波器(NSF)模型,尽管它们的波形生成方法不同,但在语音合成中表现出良好的性能。语音和音乐音频合成技术之间的相似性表明,在音乐领域应用语音合成器的最佳方式方面,探索有趣的途径。这项工作比较了三种情况下用于乐器声音生成的三种神经合成器:从音乐数据开始训练,从语音域进行零射击学习,以及从语音到音乐域的基于微调的适应。大规模感知测试的结果表明,在语音数据上进行预训练并在音乐数据上进行微调后,三种合成器的性能得到了提高,这表明语音数据的知识对音乐音频生成的有用性。在合成器中,WaveGlow在零采样学习中表现最好,而NSF在其他场景中表现最好,并且可以生成感知上接近自然音频的样本。
Recent neural waveform synthesizers such as WaveNet, WaveG-low, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similarity between speech and music audio synthesis techniques suggests interesting avenues to explore in terms of the best way to apply speech synthesizers in the music domain. This work compares three neural synthesizers used for musical instrument sounds generation under three scenarios: training from scratch on music data, zero-shot learning from the speech domain, and fine-tuning-based adaptation from the speech to the music domain. The results of a large-scale perceptual test demonstrated that the performance of three synthesizers improved when they were pre-trained on speech data and fine-tuned on music data, which indicates the usefulness of knowledge from speech data for music audio generation. Among the synthesizers, WaveGlow showed the best potential in zero-shot learning while NSF performed best in the other scenarios and could generate samples that were perceptually close to natural audio.