Word (cid:1) -gram Probability Estimation From A Japanese Raw Corpus

Word (cid:1) -gram Probability Estimation From A Japanese Raw Corpus
复制标题

来自日语原始语料库的 Word (cid:1) -gram 概率估计

DOI:
--
复制
发表时间:
2004
期刊:
--
影响因子:
--
通讯作者:
Daisuke Takuma
Daisuke Takuma
中科院分区:
--
文献类型:
--
作者:
Shinsuke Mori;Daisuke Takuma

文献摘要

被引文献

相似文献

统计语言模型在最先进的语音识别器中发挥着重要作用。最常用的语言模型(LM)是单词(cid:1)-gram模型,它基于语料库中单词和单词序列的频率。然而,在各种亚洲语言中,单词不是由空格分隔的,因此我们需要用单词边界信息来注释句子,以准备一个统计上可靠的大型语料库。在本文中,我们提出了一种直接从原始语料库构建语言模型的方法。在该方法中,原始语料库中的句子被视为带有随机词边界信息注释的句子。在实验中,我们比较了仅从分段语料库构建的 LM 和从分段语料库和原始语料库构建的 LM 的预测能力。结果表明,通过我们的方法,我们成功地将原始语料库的复杂度降低了 42.9%。
Statistical language modeling plays an important role in a state-of-the-art speech recognizer. The most used language model (LM) is word (cid:1) -gram model, which is based on the frequency of words and word sequences in a corpus. In various Asian languages, however, words are not delimited by whitespace, so we need to annotate sentences with word boundary information to prepare a statistically reliable large corpus. In this paper, we propose a method for building an LM directly from a raw corpus. In this method, sentences in the raw corpus are regarded as sentences annotated with stochastic word boundary information. In the experiments, we compared the predictive powers of an LM built only from a segmented coprus and an LM built from the segmented corpus and a raw corpus. The result showed that we succeeded in reducing the per-plexity by 42.9% using a raw corpus by our method.