Initial investigation of speech synthesis based on complex-valued neural networks

Initial investigation of speech synthesis based on complex-valued neural networks
复制标题

DOI:
10.1109/icassp.2016.7472755
复制
发表时间:
2016-03
期刊:
2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)
影响因子:
--
通讯作者:
Qiong Hu;J. Yamagishi;Korin Richmond;K. Subramanian;Y. Stylianou
Qiong Hu;J. Yamagishi;Korin Richmond;K. Subramanian;Y. Stylianou
中科院分区:
其他
文献类型:
--
作者:
Qiong Hu;J. Yamagishi;Korin Richmond;K. Subramanian;Y. Stylianou

文献摘要

被引文献

相似文献

虽然频率分析经常将我们引向复域中的语音信号,但我们经常使用的声学模型是为实值数据设计的。相位通常被忽略或与频谱幅度分开建模。在这里,我们提出了一个复值神经网络(CVNN)直接建模的结果的频率分析在复杂的域(如复振幅)。我们还引入了一种相位编码技术,将实值数据(例如,倒谱或对数振幅)映射到复域,因此我们可以无缝地使用相同的CVNN处理。在本文中,一个完全复值神经网络,即一个神经网络,其中所有的权重矩阵,激活函数和学习算法是在复域中,应用于语音合成。结果表明,它能够模拟复值和实值数据。
Although frequency analysis often leads us to a speech signal in the complex domain, the acoustic models we frequently use are designed for real-valued data. Phase is usually ignored or modelled separately from spectral amplitude. Here, we propose a complex-valued neural network (CVNN) for directly modelling the results of the frequency analysis in the complex domain (such as the complex amplitude). We also introduce a phase encoding technique to map real-valued data (e.g. cepstra or log amplitudes) into the complex domain so we can use the same CVNN processing seamlessly. In this paper, a fully complex-valued neural network, namely a neural network where all of the weight matrices, activation functions and learning algorithms are in the complex domain, is applied for speech synthesis. Results show its ability to model both complex-valued and real-valued data.