Reverberation Modeling for Source-Filter-based Neural Vocoder

Reverberation Modeling for Source-Filter-based Neural Vocoder
复制标题

DOI:
10.21437/interspeech.2020-1613
复制
发表时间:
2020-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Yang Ai;Xin Wang;J. Yamagishi;Zhenhua Ling
Yang Ai;Xin Wang;J. Yamagishi;Zhenhua Ling
中科院分区:
其他
文献类型:
--
作者:
Yang Ai;Xin Wang;J. Yamagishi;Zhenhua Ling

文献摘要

相似文献

本文提出了一种基于源滤波器的神经声编码器混响模块,提高了混响效果建模的性能。该模块使用神经声编码器的输出波形作为输入,并通过将输入与房间脉冲响应(RIR)卷积产生混响波形。我们提出了参数化和估计RIR的两种方法。第一种方法假设一个全局时不变RIR (GTI),并直接在训练数据集上学习RIR的值。第二种方法假设一个话语级时变RIR,它在一个话语中是不变的,但在不同的话语中是不同的,并使用另一个神经网络来预测RIR值。我们将所提出的混响模块加入到HiNet声码器的相谱预测器(PSP)中,并对模型进行联合训练。实验结果表明,该模型能够有效地模拟混响效应,提高混响语音的感知质量。在未知混响条件下,UTV-RIR比GTI-RIR具有更强的鲁棒性,获得了更好的混响效果。
This paper presents a reverberation module for source-filter-based neural vocoders that improves the performance of reverberant effect modeling. This module uses the output waveform of neural vocoders as an input and produces a reverberant waveform by convolving the input with a room impulse response (RIR). We propose two approaches to parameterizing and estimating the RIR. The first approach assumes a global time-invariant (GTI) RIR and directly learns the values of the RIR on a training dataset. The second approach assumes an utterance-level time-variant (UTV) RIR, which is invariant within one utterance but varies across utterances, and uses another neural network to predict the RIR values. We add the proposed reverberation module to the phase spectrum predictor (PSP) of a HiNet vocoder and jointly train the model. Experimental results demonstrate that the proposed module was helpful for modeling the reverberation effect and improving the perceived quality of generated reverberant speech. The UTV-RIR was shown to be more robust than the GTI-RIR to unknown reverberation conditions and achieved a perceptually better reverberation effect.