Large-scale random forest language models for speech recognition

Large-scale random forest language models for speech recognition
复制标题

用于语音识别的大规模随机森林语言模型

DOI:
10.21437/interspeech.2007-259
复制
发表时间:
2007
期刊:
Interspeech
影响因子:
--
通讯作者:
S. Khudanpur
S. Khudanpur
中科院分区:
--
文献类型:
--
作者:
Yi Su;F. Jelinek;S. Khudanpur

文献摘要

被引文献

相似文献

随机森林语言模型(RFLM)在几个自动语音识别(ASR)任务中表现出令人鼓舞的结果,但受到实际限制的阻碍,特别是从大量数据中估计RFLM的空间复杂性。本文通过一种有效的磁盘交换策略来解决RFLM的大规模训练和测试问题,该策略利用了二叉决策树的递归结构和树生长算法的局部访问特性,从而充分发挥了RFLM的潜力,并开辟了进一步研究的途径,包括与n元模型的有用比较。通过使用最先进的ASR系统的困惑减少和晶格重新评分实验,证明了该策略的贝内拟合。
The random forest language model (RFLM) has shown encouraging results in several automatic speech recognition (ASR) tasks but has been hindered by practical limitations, notably the space-complexity of RFLM estimation from large amounts of data. This paper addresses large-scale training and testing of the RFLM via an efficient disk-swapping strategy that exploits the recursive structure of a binary decision tree and the local access property of the tree-growing algorithm, redeeming the full potential of the RFLM, and opening avenues of further research, including useful comparisons with n -gram models. Benefits of this strategy are demonstrated by perplexity reduction and lattice rescoring experiments using a state-of-the-art ASR system.