Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASR

Improving OOV Detection and Resolution with External Language Models in Acoustic-to-Word ASR
复制标题

DOI:
10.1109/slt.2018.8639563
复制
发表时间:
2018-12
期刊:
2018 IEEE Spoken Language Technology Workshop (SLT)
影响因子:
--
通讯作者:
H. Inaguma;M. Mimura;S. Sakai;Tatsuya Kawahara
H. Inaguma;M. Mimura;S. Sakai;Tatsuya Kawahara
中科院分区:
其他
文献类型:
--
作者:
H. Inaguma;M. Mimura;S. Sakai;Tatsuya Kawahara

文献摘要

相似文献

声学到单词(A2W)端到端自动语音识别(ASR)系统因其极其简化的体系结构和快速的解码而受到关注。为了缓解由于不常用单词而导致的数据稀疏问题,研究了与声学到字符(A2C)模型的结合。此外,A2C模型可以用于恢复A2W模型未涵盖的词汇表(OOV)单词,但这需要准确检测OOV单词。A2W模型通过声学和抄本学习上下文;因此,他们倾向于错误地将OOV单词识别为词汇中的单词。在本文中,我们通过使用外部语言模型(LM)来解决这个问题,该模型只用转录进行训练,并且具有更好的语言信息来检测OOV词。A2C模型被用来解析这些OOV词。实验结果表明,外部LMS不仅可以减少错误,而且可以增加OOV词数,在英语会话和日语演讲语料库中的性能得到显著提高,尤其是在领域外场景中。我们还研究了A2W模型的词汇量和训练LMS的数据量的影响。此外,我们的方法可以在性能略微下降的情况下将词汇量减少几倍。
Acoustic-to-word (A2W) end-to-end automatic speech recognition (ASR) systems have attracted attention because of an extremely simplified architecture and fast decoding. To alleviate data sparseness issues due to infrequent words, the combination with an acoustic-to-character (A2C) model is investigated. Moreover, the A2C model can be used to recover-of-vocabulary (OOV) words that are not covered by the A2W model, but this requires accurate detection of OOV words. A2W models learn contexts with both acoustic and transcripts; therefore they tend to falsely recognize OOV words as words in the vocabulary. In this paper, we tackle this problem by using external language models (LM), which are trained only with transcriptions and have better linguistic information to detect OOV words. The A2C model is used to resolve these OOV words. Experimental evaluations show that external LMs have the effects of not only reducing errors but also increasing the number of detected OOV words, and the proposed method significantly improves performances in English conversational and Japanese lecture corpora, especially for-of-domain scenario. We also investigate the impact of the vocabulary size of A2W models and the data size for training LMs. Moreover, our approach can reduce the vocabulary size several times with marginal performance degradation.