Training a Language Model Using Webdata for Large Vocabulary Japanese Spontaneous Speech Recognition

Training a Language Model Using Webdata for Large Vocabulary Japanese Spontaneous Speech Recognition
复制标题

使用网络数据训练语言模型进行大词汇日语自发语音识别

DOI:
10.21437/interspeech.2011-258
复制
发表时间:
2011
期刊:
--
影响因子:
--
通讯作者:
Akinori Ito
Akinori Ito
中科院分区:
--
文献类型:
--
作者:
Ryo Masumura;Seongjun Hahm;Akinori Ito

文献摘要

被引文献

相似文献

本文描述了一种语言建模方法,使用大规模的口语数据从Web上检索的自发语音识别。我们下载了1500万个关于广泛主题的网页。接下来,使用朴素贝叶斯分类器从下载的Web数据中选择类似口语的文本,并使用模拟模型添加典型的语言现象,如莱尔和停顿。由生成的数据训练的语言模型给出了与大规模自发语音语料库(语料库自发日语,CSJ)一样高的性能。通过结合生成的数据和CSJ,我们提高了单词准确性。
This paper describes a language modeling method using large-scale spoken language data retrieved from the Web for spon-taneous speech recognition. We downloaded 15 million Web pages on a comprehensive range topics. Next, spoken language-like texts were selected from the downloaded Web data using the na¨ıve Bayes classifier, and typical linguistic phenomena such as fillers and pauses were added using simulation models. A language model trained by the generated data gave as high performance as the large-scale spontaneous speech corpus (Corpus of Spontaneous Japanese, CSJ). By combining the generated data and CSJ, we improved word accuracy.