LSTBM: A Novel Sequence Representation of Speech Spectra Using Restricted Boltzmann Machine with Long Short-Term Memory

LSTBM: A Novel Sequence Representation of Speech Spectra Using Restricted Boltzmann Machine with Long Short-Term Memory
复制标题

DOI:
10.21437/interspeech.2018-1753
复制
发表时间:
2018-09
期刊:
--
影响因子:
--
通讯作者:
Toru Nakashika
Toru Nakashika
中科院分区:
其他
文献类型:
--
作者:
Toru Nakashika

文献摘要

相似文献

在本文中,我们提出了一种新的概率模型,即长短期玻尔兹曼记忆(LSTBM),表示序列数据,如语音频谱。LSTBM是受限玻尔兹曼机(RBM)的扩展,它具有生成的长短期记忆(LSTM)单元。原始的RBM自动学习可见和隐藏单元之间的关系,并被广泛用作特征提取器,生成器,分类器,深度神经网络的预训练方法等。与传统的RBM不同,LSTBM通过LSTM单元随着时间的推移而连接,并表示顺序数据中的时间依赖性。我们的语音编码实验表明,所提出的LSTBM优于其他传统的方法:RBM和时间RBM。
In this paper, we propose a novel probabilistic model, namely long short-term Boltzmann memory (LSTBM), to represent sequential data like speech spectra. The LSTBM is an extension of a restricted Boltzmann machine (RBM) that has generative long short-term memory (LSTM) units. The original RBM automatically learns relationships between visible and hidden units and is widely used as a feature extractor, a generator, a classifier, a pre-training method of deep neural networks, etc. However, the RBM is not sufficient to represent sequential data because it assumes that each frame from sequential data is completely independent of the others. Unlike conventional RBMs, the LSTBM has connections over time via LSTM units and represents time dependencies in sequential data. Our speech coding experiments demonstrated that the proposed LSTBM out-performed the other conventional methods: an RBM and a temporal RBM.