Building topic specific language models from webdata using competitive models

Building topic specific language models from webdata using competitive models
复制标题

使用竞争模型从网络数据构建主题特定语言模型

DOI:
10.21437/interspeech.2005-20
复制
发表时间:
2005
期刊:
--
影响因子:
--
通讯作者:
Shrikanth S. Narayanan
Shrikanth S. Narayanan
中科院分区:
--
文献类型:
--
作者:
A. Sethy;P. Georgiou;Shrikanth S. Narayanan

文献摘要

被引文献

相似文献

快速构建主题规范fic语言模型的能力是ASR跨不同领域快速部署和可移植性的关键需求。万维网有望成为创建主题规范fic语言模型的极好的文本数据资源。本文描述了一种迭代网络爬行方法,它使用一组相互竞争的自适应模型,包括通用的主题无关背景语言模型、表示基于Web的数据(WebData)中遇到的虚假文本的噪声模型和主题特定fic模型,以使用基于相对熵的方法为万维网搜索引擎生成查询串,并对下载的Web数据进行适当的加权,以建立主题特定fic语言模型。我们演示了如何使用这个系统来快速构建特定fic领域的语言模型,只给出了一组初始的示例话语,以及它如何解决与Webdata相关的各种问题。在我们的实验中,我们能够为我们的目标医疗领域减少20%的困惑。困惑方面的收益转化为ASR单词错误率(绝对值)4%的改善,相应地相对收益为14%。
The ability to build topic specific language models, rapidly and with minimal human effort, is a critical need for fast deployment and portability of ASR across different domains. The World Wide Web (WWW) promises to be an excellent textual data resource for creating topic specific language models. In this paper we describe an iterative web crawling approach which uses a competitive set of adaptive models comprised of a generic topic independent background language model, a noise model representing spurious text encountered in web based data (Webdata), and a topic specific model to generate query strings using a relative entropy based approach for WWW search engines and to weight the downloaded Webdata appropriately for building topic specific language models. We demonstrate how this system can be used to rapidly build language models for a specific domain given just an initial set of example utterances and how it can address the various issues attached with Webdata. In our experiments we were able to achieve a 20% reduction in perplexity for our target medical domain. The gains in perplexity translated to a 4% improvement in ASR word error rate (absolute) corresponding to a relative gain of 14%.