PeriodNet: A Non-Autoregressive Raw Waveform Generative Model With a Structure Separating Periodic and Aperiodic Components

PeriodNet: A Non-Autoregressive Raw Waveform Generative Model With a Structure Separating Periodic and Aperiodic Components
复制标题

DOI:
10.1109/access.2021.3118033
复制
发表时间:
2021
期刊:
影响因子:
3.9
通讯作者:
Yukiya Hono;Shinji Takaki;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda
Yukiya Hono;Shinji Takaki;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yukiya Hono;Shinji Takaki;Kei Hashimoto;Keiichiro Oura;Yoshihiko Nankaku;K. Tokuda

文献摘要

相似文献

本文提出了PeriodNet,一种非自回归(非AR)波形生成模型,具有新的模型结构,用于对语音波形中的周期性和非周期性分量进行建模。非 AR 原始波形生成模型可以快速生成高质量波形。然而,这些模型可以重建的波形变化受到训练数据的限制。此外,典型的非 AR 模型从单个高斯输入重建语音波形,尽管语音中混合了周期性和非周期性信号。这些可能会显着影响某些应用中的波形生成过程,例如歌声合成系统,这些应用需要以较少的周期性再现准确的音高和自然声音,包括哈士奇声和呼吸声。 periodNet 使用并行或串行模型结构对语音波形进行建模来解决这些问题。并联或串联的两个子发生器将显式周期和非周期信号(正弦波和高斯噪声)作为输入。由于PeriodNet通过关注这些输入信号是否自相关来对周期性和非周期性分量进行建模,因此在训练期间不需要外部周期性/非周期性分解。实验结果表明,我们提出的结构提高了生成波形的自然度。我们还表明,可以更自然地生成音调超出训练数据范围的语音波形。
This paper presents PeriodNet, a non-autoregressive (non-AR) waveform generative model with a new model structure for modeling periodic and aperiodic components in speech waveforms. Non-AR raw waveform generative models have enabled the fast generation of high-quality waveforms. However, the variations of waveforms that these models can reconstruct are limited by training data. In addition, typical non-AR models reconstruct a speech waveform from a single Gaussian input despite the mixture of periodic and aperiodic signals in speech. These may significantly affect the waveform generation process in some applications such as singing voice synthesis systems, which require reproducing accurate pitch and natural sounds with less periodicity, including husky and breath sounds. PeriodNet uses a parallel or series model structure to model a speech waveform to tackle these problems. Two sub-generators connected in parallel or in series take an explicit periodic and aperiodic signal (sine wave and Gaussian noise) as an input. Since PeriodNet models periodic and aperiodic components by focusing on whether these input signals are autocorrelated or not, it does not require external periodic/aperiodic decomposition during training. Experimental results show that our proposed structure improves the naturalness of generated waveforms. We also show that speech waveforms with a pitch outside of the training data range can be generated with more naturalness.