A Part of Speech Estimation Method for Japanese Unknown Words using a Statistical Model of Morphology and Context

A Part of Speech Estimation Method for Japanese Unknown Words using a Statistical Model of Morphology and Context
复制标题

使用词法和语境统计模型的日语未知词的部分语音估计方法

DOI:
10.3115/1034678.1034725
复制
发表时间:
1999
期刊:
--
影响因子:
--
通讯作者:
M. Nagata
M. Nagata
中科院分区:
--
文献类型:
--
作者:
M. Nagata

文献摘要

被引文献

相似文献

我们提出了一个日语未知词的统计模型,该模型由一组长度和拼写模型组成,这些模型由组成单词的字符类型分类。重点很简单:不同的字符集应该被区别对待,字符类型之间的变化非常重要,因为日文既有汉字(汉字)这样的表意文字,也有英语(片假名)这样的声像文字。该模型提高了分词精度和词性标注精度。在正确分割未知词的情况下,该模型的标注准确率可达96.6%。
We present a statistical model of Japanese unknown words consisting of a set of length and spelling models classified by the character types that constitute a word. The point is quite simple: different character sets should be treated differently and the changes between character types are very important because Japanese script has both ideograms like Chinese (kanji) and phonograms like English (katakana). Both word segmentation accuracy and part of speech tagging accuracy are improved by the proposed model. The model can achieve 96.6% tagging accuracy if unknown words are correctly segmented.