A Part of Speech Estimation Method for Japanese Unknown Words using a Statistical Model of Morphology and Context
A Part of Speech Estimation Method for Japanese Unknown Words using a Statistical Model of Morphology and Context
复制标题
使用词法和语境统计模型的日语未知词的部分语音估计方法
DOI:
10.3115/1034678.1034725
复制
发表时间:
1999
期刊:
影响因子:
--
通讯作者:
M. Nagata
中科院分区:
文献类型:
--
作者:
M. Nagata
We present a statistical model of Japanese unknown words consisting of a set of length and spelling models classified by the character types that constitute a word. The point is quite simple: different character sets should be treated differently and the changes between character types are very important because Japanese script has both ideograms like Chinese (kanji) and phonograms like English (katakana). Both word segmentation accuracy and part of speech tagging accuracy are improved by the proposed model. The model can achieve 96.6% tagging accuracy if unknown words are correctly segmented.