Language model adaptation using WWW documents obtained by utterance-based queries
Language model adaptation using WWW documents obtained by utterance-based queries
复制标题
使用基于话语的查询获得的 WWW 文档进行语言模型自适应
DOI:
10.1109/icassp.2010.5494928
复制
发表时间:
2010
期刊:
影响因子:
--
通讯作者:
Shrikanth S. Narayanan
中科院分区:
文献类型:
--
作者:
A. Tsiartas;P. Georgiou;Shrikanth S. Narayanan
In this paper, we consider the estimation of topic specific Language Models (LM) by exploiting documents from the World Wide Web (WWW). We focus on the quality of the generated queries and propose a novel query generation method. In contrast to the n-gram based queries used in past works, our approach relies on utterances as queries candidates. The proposed approach does not rely on any language specific information other than the initial in-domain training text. We have conducted experiments with Web texts of size 0–150 million words, and we have shown that despite not using any language specific information, the proposed approach results in up to 1.1% absolute Word Error Rate (WER) improvement as compared to keyword-based approaches. The proposed approach reduces the WER by 6.3% absolute in our experiments, compared to an in-domain LM without considering any Web data.