Probabilistic Modeling of Korean Morphology

Probabilistic Modeling of Korean Morphology
复制标题

韩国语形态学的概率建模

DOI:
10.1109/tasl.2009.2019922
复制
发表时间:
2009
期刊:
IEEE Transactions on Audio, Speech, and Language Processing
影响因子:
--
通讯作者:
Hae
Hae
中科院分区:
--
文献类型:
--
作者:
Do;Hae

文献摘要

被引文献

相似文献

本文提出了一种新的朝鲜语词法分析概率模型。为了利用韩国语词法的特点,所提出的模型基于三个语言单位:eojeol(韩国语的间距单位),语素和音节。与以往基于规则和词典的方法不同,本文提出的概率方法可以自动从词性标注的语料库中获取完整的语言知识。此外,这种方法无需任何系统修改,可以很容易地适用于具有不同标记集和注释指南的其他语料库。这三种不同的模型及其组合在三个语料库上进行了广泛的条件评估。词性单位和音节单位模型弥补了语素单位模型的不足。eojeol-unit模型运行效率高,精度高。音节单元模型的精度也得到了提高,在处理未知单词方面表现出了特别强劲的表现。所提出的方法也被证明优于先前的方法。
This paper proposes new probabilistic models for analyzing Korean morphology. In order to take advantage of the characteristics of Korean morphology, the proposed models are based on three linguistic units: eojeol (a Korean spacing unit), morpheme, and syllable. Unlike previous approaches that are based on rules and dictionaries, the probabilistic approach proposed in this study can automatically acquire complete linguistic knowledge from part-of-speech (POS) tagged corpora. In addition, this approach, without any system modification, is easily applicable to other corpora with different tag sets and annotation guidelines. The three different models and their combinations are evaluated on three corpora over a wide range of conditions. The eojeol-unit and syllable-unit models compensate for the weaknesses of the morpheme-unit model. The eojeol-unit model performed efficiently, and improved the precision. The syllable-unit model improved in precision as well, showing a particularly robust performance in treating unknown words. The proposed approach is also proven to outperform the previous approaches.