Discriminative Reranking for LVCSR Leveraging Invariant Structure
Discriminative Reranking for LVCSR Leveraging Invariant Structure
复制标题
DOI:
10.21437/interspeech.2012-173
复制
发表时间:
2012
期刊:
影响因子:
4.5
通讯作者:
Masayuki Suzuki;Gakuto Kurata;M. Nishimura;N. Minematsu
中科院分区:
文献类型:
--
作者:
Masayuki Suzuki;Gakuto Kurata;M. Nishimura;N. Minematsu
An invariant structure is one of the long-span acoustic representations, where acoustic variations caused by non-linguistic factors are effectively removed from speech. We present in this paper a new method to leverage the invariant structures as features of discriminative reranking for Large Vocabulary Continuous Speech Recognition (LVCSR). First we use a traditional HMMbased LVCSR system to get a list of N -best candidates with phone alignments and construct an invariant structure for each candidate using its phone alignment. Here, the invariant structure is composed of lengths between every two phonemes in the candidate. Then we estimate a score of each phoneme-pair in the invariant structure, and rerank the N -best candidates using a weighted sum of the phoneme-pair scores, where the weights are trained discriminatively by averaged perceptron. Experimental results show a relative CER improvement of 6.69% over the baseline HMM-based LVCSR system.