Hidden Conditional Neural Fields for Continuous Phoneme Speech Recognition

Hidden Conditional Neural Fields for Continuous Phoneme Speech Recognition
复制标题

DOI:
10.1587/transinf.e95.d.2094
复制
发表时间:
2012-08
期刊:
IEICE Trans. Inf. Syst.
影响因子:
--
通讯作者:
Yasuhisa Fujii;Kazumasa Yamamoto;S. Nakagawa
Yasuhisa Fujii;Kazumasa Yamamoto;S. Nakagawa
中科院分区:
其他
文献类型:
--
作者:
Yasuhisa Fujii;Kazumasa Yamamoto;S. Nakagawa

文献摘要

相似文献

摘要:在本文中,我们提出了用于连续音素语音识别的隐藏条件神经场(HCNF),它是隐藏条件随机场(HCRF)和多层感知器(MLP)的组合,并继承了它们的优点,即 HCRF 对序列的判别性和从 MLP 中提取非线性特征的能力。 HCNF 可以合并多种类型的特征,从中提取非线性特征,并通过顺序标准进行训练。我们首先提出 HCNF 的公式,然后研究使用 HCNF 进一步改进自动语音识别的三种方法,HCNF 是一个明确考虑训练误差的目标函数,提供分层串联式特征,并包括用于观察函数的深度非线性特征提取器。我们使用 TIMIT 核心测试集上的连续英语音素识别和日语音素识别的实验结果表明,HCNF 可以在没有任何初始模型的情况下进行实际训练,并且优于 HCRF 和通过最小音素误差 (MPE) 方式训练的三音素隐马尔可夫模型
SUMMARY In this paper, we propose Hidden Conditional Neural Fields (HCNF) for continuous phoneme speech recognition, which are a combination of Hidden Conditional Random Fields (HCRF) and a MultiLayer Perceptron (MLP), and inherit their merits, namely, the discriminative property for sequences from HCRF and the ability to extract non-linear features from an MLP. HCNF can incorporate many types of features from which non-linear features can be extracted, and is trained by sequential criteria. We first present the formulation of HCNF and then examine three methods to further improve automatic speech recognition using HCNF, which is an objective function that explicitly considers training errors, provides a hierarchical tandem-style feature and includes a deep non-linear feature extractor for the observation function. We show that HCNF can be trained realistically without any initial model and outperforms HCRF and the triphone hidden Markov model trained by the minimum phone error (MPE) manner using experimental results for continuous English phoneme recognition on the TIMIT core test set and Japanese phoneme recognition