Complex-Valued Restricted Boltzmann Machine for Speaker-Dependent Speech Parameterization From Complex Spectra

Complex-Valued Restricted Boltzmann Machine for Speaker-Dependent Speech Parameterization From Complex Spectra
复制标题

DOI:
10.1109/taslp.2018.2877465
复制
发表时间:
2019-02
期刊:
IEEE/ACM Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Toru Nakashika;Shinji Takaki;J. Yamagishi
Toru Nakashika;Shinji Takaki;J. Yamagishi
中科院分区:
其他
文献类型:
--
作者:
Toru Nakashika;Shinji Takaki;J. Yamagishi

文献摘要

相似文献

本文描述了一种新的基于能量的概率分布,代表复值数据,并解释了如何将其应用于直接从复值谱特征提取。所提出的模型,复值限制玻尔兹曼机(CRBM),旨在处理复值可见单位作为一个扩展的著名的限制玻尔兹曼机(RBM)。与RBM一样,CRBM学习可见和隐藏单元之间的关系,而不需要同一层中单元之间的连接,这通过使用吉布斯采样或对比发散来显着提高训练效率。另一个重要的特征是CRBM还具有每个复值可见单元的真实的部分和虚部之间的连接,这有助于表示复域中的数据分布。在语音信号处理中,分类和生成特征通常基于幅度谱(例如,MFCC、倒谱和梅尔倒谱),即使它们是从复谱计算的,并且它们忽略相位信息。相比之下,所提出的特征提取器使用CRBM直接编码的复谱(或另一个复值表示的复谱)到二进制值的潜在特征(隐藏单元)。由于可见-隐藏连接是无向的,我们还可以直接从潜在特征中恢复(解码)复杂的光谱。我们的语音表示实验表明,CRBM优于其他语音表示方法,如使用传统的RBM,梅尔对数谱近似解码器等方法。
This paper describes a novel energy-based probabilistic distribution that represents complex-valued data and explains how to apply it to direct feature extraction from complex-valued spectra. The proposed model, the complex-valued restricted Boltzmann machine (CRBM), is designed to deal with complex-valued visible units as an extension of the well-known restricted Boltzmann machine (RBM). Like the RBM, the CRBM learns the relationships between visible and hidden units without having connections between units in the same layer, which dramatically improves training efficiency by using Gibbs sampling or contrastive divergence. Another important characteristic is that the CRBM also has connections between real and imaginary parts of each of the complex-valued visible units that help represent the data distribution in the complex domain. In speech signal processing, classification and generation features are often based on amplitude spectra (e.g., MFCC, cepstra, and mel-cepstra) even if they are calculated from complex spectra, and they ignore phase information. In contrast, the proposed feature extractor using the CRBM directly encodes the complex spectra (or another complex-valued representation of the complex spectra) into binary-valued latent features (hidden units). Since the visible-hidden connections are undirected, we can also recover (decode) the complex spectra from the latent features directly. Our speech representation experiments demonstrated that the CRBM outperformed other speech representation methods, such as methods using a conventional RBM, a mel-log spectrum approximate decoder, etc.