Language model adaptation using WWW documents obtained by utterance-based queries

Language model adaptation using WWW documents obtained by utterance-based queries
复制标题

使用基于话语的查询获得的 WWW 文档进行语言模型自适应

DOI:
10.1109/icassp.2010.5494928
复制
发表时间:
2010
期刊:
2010 IEEE International Conference on Acoustics, Speech and Signal Processing
影响因子:
--
通讯作者:
Shrikanth S. Narayanan
Shrikanth S. Narayanan
中科院分区:
--
文献类型:
--
作者:
A. Tsiartas;P. Georgiou;Shrikanth S. Narayanan

文献摘要

被引文献

相似文献

在本文中,我们考虑估计的主题特定的语言模型(LM),利用文件从万维网(WWW)。我们专注于生成的查询的质量,并提出了一种新的查询生成方法。与过去的作品中使用的基于n-gram的查询相比,我们的方法依赖于话语作为查询候选。所提出的方法不依赖于任何语言特定的信息,而不是最初的域内训练文本。我们已经进行了实验与Web文本的大小0- 1.5亿字,我们已经表明,尽管不使用任何语言的特定信息,所提出的方法的结果在高达1.1%的绝对字错误率(WER)的改善相比,基于关键字的方法。在我们的实验中,所提出的方法减少了6.3%的绝对WER,相比,在域LM不考虑任何Web数据。
In this paper, we consider the estimation of topic specific Language Models (LM) by exploiting documents from the World Wide Web (WWW). We focus on the quality of the generated queries and propose a novel query generation method. In contrast to the n-gram based queries used in past works, our approach relies on utterances as queries candidates. The proposed approach does not rely on any language specific information other than the initial in-domain training text. We have conducted experiments with Web texts of size 0–150 million words, and we have shown that despite not using any language specific information, the proposed approach results in up to 1.1% absolute Word Error Rate (WER) improvement as compared to keyword-based approaches. The proposed approach reduces the WER by 6.3% absolute in our experiments, compared to an in-domain LM without considering any Web data.