Word (cid:1) -gram Probability Estimation From A Japanese Raw Corpus
Word (cid:1) -gram Probability Estimation From A Japanese Raw Corpus
复制标题
来自日语原始语料库的 Word (cid:1) -gram 概率估计
DOI:
--
复制
发表时间:
2004
期刊:
影响因子:
--
通讯作者:
Daisuke Takuma
中科院分区:
文献类型:
--
作者:
Shinsuke Mori;Daisuke Takuma
Statistical language modeling plays an important role in a state-of-the-art speech recognizer. The most used language model (LM) is word (cid:1) -gram model, which is based on the frequency of words and word sequences in a corpus. In various Asian languages, however, words are not delimited by whitespace, so we need to annotate sentences with word boundary information to prepare a statistically reliable large corpus. In this paper, we propose a method for building an LM directly from a raw corpus. In this method, sentences in the raw corpus are regarded as sentences annotated with stochastic word boundary information. In the experiments, we compared the predictive powers of an LM built only from a segmented coprus and an LM built from the segmented corpus and a raw corpus. The result showed that we succeeded in reducing the per-plexity by 42.9% using a raw corpus by our method.