Ergodic multigram HMM integrating word segmentation and class tagging for Chinese language modeling

Ergodic multigram HMM integrating word segmentation and class tagging for Chinese language modeling
复制标题

DOI:
10.1109/icassp.1996.540324
复制
发表时间:
1996-05
期刊:
1996 IEEE International Conference on Acoustics, Speech, and Signal Processing Conference Proceedings
影响因子:
--
通讯作者:
Hubert Hin-Cheung Law;Chorkin Chan
Hubert Hin-Cheung Law;Chorkin Chan
中科院分区:
其他
文献类型:
--
作者:
Hubert Hin-Cheung Law;Chorkin Chan

文献摘要

被引文献

相似文献

提出了一种新的遍历多字组隐马尔可夫模型(HMM),该模型将句子生成建模为一个双随机过程,首先根据一阶马尔可夫模型生成词类,然后基于词类独立生成单字词或多字词,句子上不标注词边界.该模型适用于汉语等没有词边界标记的语言。它的词典包含每个单词的语法类,其应用包括识别器的语言建模,以及集成的分词和类标记。训练时不需要预先分割和标记的语料,分割和标记都在一个模型中训练。本文给出了该模型的相关算法,并在中文新闻语料库上进行了实验。
A novel ergodic multigram hidden Markov model (HMM) is introduced which models sentence production as a doubly stochastic process, in which word classes are first produced according to a first order Markov model, and then single or multi-character words are generated independently based on the word classes, without word boundary marked on the sentence. This model can be applied to languages without word boundary markers such as Chinese. With a lexicon containing syntactic classes for each word, its applications include language modeling for recognizers, and integrated word segmentation and class tagging. Pre-segmented and tagged corpus are not needed for training, and both segmentation and tagging are trained in one single model. In this paper, relevant algorithms for this model are presented, and experimental results on a Chinese news corpus are reported.