Discriminative Reranking for LVCSR Leveraging Invariant Structure

Discriminative Reranking for LVCSR Leveraging Invariant Structure
复制标题

DOI:
10.21437/interspeech.2012-173
复制
发表时间:
2012
期刊:
影响因子:
4.5
通讯作者:
Masayuki Suzuki;Gakuto Kurata;M. Nishimura;N. Minematsu
Masayuki Suzuki;Gakuto Kurata;M. Nishimura;N. Minematsu
中科院分区:
医学3区
文献类型:
--
作者:
Masayuki Suzuki;Gakuto Kurata;M. Nishimura;N. Minematsu

文献摘要

相似文献

不变结构是大跨度声学表示之一,其中由非语言因素引起的声学变化被有效地从语音中消除。我们在本文中提出了一种新方法,利用不变结构作为大词汇量连续语音识别(LVCSR)的判别性重排序特征。首先,我们使用传统的基于 HMM 的 LVCSR 系统来获取具有音素对齐的 N 个最佳候选列表,并使用其音素对齐为每个候选构建一个不变结构。这里,不变结构由候选中每两个音素之间的长度组成。然后,我们估计不变结构中每个音素对的分数,并使用音素对分数的加权和对 N 个最佳候选重新排序,其中权重由平均感知器进行有区别的训练。实验结果表明,与基于 HMM 的基线 LVCSR 系统相比,相对 CER 提高了 6.69%。
An invariant structure is one of the long-span acoustic representations, where acoustic variations caused by non-linguistic factors are effectively removed from speech. We present in this paper a new method to leverage the invariant structures as features of discriminative reranking for Large Vocabulary Continuous Speech Recognition (LVCSR). First we use a traditional HMMbased LVCSR system to get a list of N -best candidates with phone alignments and construct an invariant structure for each candidate using its phone alignment. Here, the invariant structure is composed of lengths between every two phonemes in the candidate. Then we estimate a score of each phoneme-pair in the invariant structure, and rerank the N -best candidates using a weighted sum of the phoneme-pair scores, where the weights are trained discriminatively by averaged perceptron. Experimental results show a relative CER improvement of 6.69% over the baseline HMM-based LVCSR system.