Training a Language Model Using Webdata for Large Vocabulary Japanese Spontaneous Speech Recognition
Training a Language Model Using Webdata for Large Vocabulary Japanese Spontaneous Speech Recognition
复制标题
使用网络数据训练语言模型进行大词汇日语自发语音识别
DOI:
10.21437/interspeech.2011-258
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
Akinori Ito
中科院分区:
文献类型:
--
作者:
Ryo Masumura;Seongjun Hahm;Akinori Ito
This paper describes a language modeling method using large-scale spoken language data retrieved from the Web for spon-taneous speech recognition. We downloaded 15 million Web pages on a comprehensive range topics. Next, spoken language-like texts were selected from the downloaded Web data using the na¨ıve Bayes classifier, and typical linguistic phenomena such as fillers and pauses were added using simulation models. A language model trained by the generated data gave as high performance as the large-scale spontaneous speech corpus (Corpus of Spontaneous Japanese, CSJ). By combining the generated data and CSJ, we improved word accuracy.