Rapid transition to new spoken dialogue domains: language model training using knowledge from previous domain applications and web text resources

Rapid transition to new spoken dialogue domains: language model training using knowledge from previous domain applications and web text resources
复制标题

快速过渡到新的口语对话领域:使用先前领域应用程序和网络文本资源的知识进行语言模型训练

DOI:
10.21437/interspeech.2005-590
复制
发表时间:
2005
期刊:
--
影响因子:
--
通讯作者:
H. Kuo
H. Kuo
中科院分区:
--
文献类型:
--
作者:
Murat Akbacak;Yuqing Gao;L. Gu;H. Kuo

文献摘要

被引文献

相似文献

在一般的自动语音识别(ASR)系统中,语言模型(LMS)通常被训练成在广泛的输入条件范围内工作。在特定于域的口语对话系统(SDSS)中使用的ASR系统在内容和风格方面受到更多限制。训练和操作条件之间的内容和/或风格的不匹配导致对话应用的性能降低。本文的主要重点是开发工具,通过自动收集文本数据来促进在语言模型训练的背景下快速开发口语对话应用程序,该文本数据有助于为新的目标领域训练准确的语言模型,而不需要手动收集任何域内数据。我们研究了一个从以前的域和万维网(WWW)中提取有用信息的框架。我们通过向搜索引擎提交查询来收集数据,然后通过语法和语义过滤来清理结果文本。接下来是人工句子的生成。在不使用任何领域内数据的情况下,我们的系统获得了19.33%的单词错误率,这一性能与人工收集的32K领域内句子的语言模型的性能相当。使用不到1%的领域内数据和自动生成的文本,我们的系统获得了接近60K领域内句子训练的语言模型的ASR性能。
In generic automatic speech recognition (ASR) systems, typically, language models (LMs) are trained to work within a broad range of input conditions. ASR systems used in domain-specific spoken dialogue systems (SDSs) are more constrained in terms of content and style. A mismatch in content and/or style between training and operating conditions results in performance degradation for the dialogue application. The main focus of this paper is to develop tools to facilitate rapid development of spoken dialogue applications within the context of language model training by focusing on the problem of automatically collecting text data that is useful to train accurate language models for the new target domain without manually collecting any in-domain data. We investigate a framework to extract useful information from previous domains and World Wide Web (WWW). We collect data by submitting queries to a search engine and then clean the resulting text via syntactic and semantic filtering. This is followed by artificial sentence generation. Without using any in-domain data, our system achieved a word error rate (WER) of 19.33%, a performance comparable to that achieved by a language model trained on manually collected 32K in-domain sentences. Using less than 1% of in-domain data along with the automatically generated text, our system achieved an ASR performance close to a language model trained on 60K in-domain sentences.