Using Random Forests in the Structured Language Model

Using Random Forests in the Structured Language Model
复制标题

DOI:
--
复制
发表时间:
2004-12
期刊:
--
影响因子:
--
通讯作者:
P. Xu;F. Jelinek
P. Xu;F. Jelinek
中科院分区:
其他
文献类型:
--
作者:
P. Xu;F. Jelinek

文献摘要

被引文献

相似文献

在本文中,我们探讨了随机森林(RFs)在结构化语言模型(SLM)中的使用,该模型使用丰富的句法信息来预测下一个单词。在这项工作中的目标是构建随机增长的决策树(DT)使用句法信息的RF和调查的性能的SLM建模的RF在自动语音识别。最初作为分类器开发的RF是决策树分类器的组合。每棵树都是基于独立采样的随机训练数据生长的,并且对于森林中的所有树都具有相同的分布,并且在决策树的每个节点上随机选择可能的问题。我们的方法扩展了RFs的原始思想,以处理语言建模中遇到的数据稀疏问题。RF已经在n-gram语言建模的背景下进行了研究,并且已经被证明可以很好地推广到看不见的数据。在本文中,我们表明,使用语法信息的RF也可以实现更好的性能,在困惑(PPL)和单词错误率(WER)在一个大词汇量的语音识别系统,相比,使用Kneser-Ney平滑的基线。
In this paper, we explore the use of Random Forests (RFs) in the structured language model (SLM), which uses rich syntactic information in predicting the next word based on words already seen. The goal in this work is to construct RFs by randomly growing Decision Trees (DTs) using syntactic information and investigate the performance of the SLM modeled by the RFs in automatic speech recognition. RFs, which were originally developed as classifiers, are a combination of decision tree classifiers. Each tree is grown based on random training data sampled independently and with the same distribution for all trees in the forest, and a random selection of possible questions at each node of the decision tree. Our approach extends the original idea of RFs to deal with the data sparseness problem encountered in language modeling. RFs have been studied in the context of n-gram language modeling and have been shown to generalize well to unseen data. We show in this paper that RFs using syntactic information can also achieve better performance in both perplexity (PPL) and word error rate (WER) in a large vocabulary speech recognition system, compared to a baseline that uses Kneser-Ney smoothing.