Automatic Annotation for Korean--Approach Based on the Contextual Exploration Method

Automatic Annotation for Korean--Approach Based on the Contextual Exploration Method
复制标题

DOI:
10.1109/dexa.2007.62
复制
发表时间:
2007-09
期刊:
18th International Workshop on Database and Expert Systems Applications (DEXA 2007)
影响因子:
--
通讯作者:
Hyunzoo Chai
Hyunzoo Chai
中科院分区:
其他
文献类型:
--
作者:
Hyunzoo Chai

文献摘要

被引文献

相似文献

我们提出了一个基于上下文探索方法的韩语自动语义标注系统。为韩语创建一个词法分析器和词性标注器是很困难的,因为它是一种高度粘着的语言。因此,按照与屈折语言相同的顺序处理韩语--形态分析,然后是句法分析,然后是语义分析--并没有产生令人满意的结果。我们的新方法识别语义信息的韩语文本,而不需要通过形态和句法分析步骤。我们的初始系统正确地注释了大约88%的标准韩语句子,并且这个注释率在文本域中保持不变。在此之前,上下文探索方法已经成功地应用于法语和阿拉伯语等多种语言。鉴于我们在韩语上的成功,我们相信这种方法可以应用于其他黏着语言,如日语,土耳其语和芬兰语。
We present an automatic semantic annotation system for Korean based on the contextual exploration method. Creating a morphological analyzer and part-of-speech tagger for the Korean language is difficult as it is a highly agglutinative language. Accordingly, processing Korean in the same order as inflectional languages - morphological analysis, then syntactical and then semantic - has not yielded satisfactory results. Our new method identifies semantic information in Korean text without going through the morphological and syntactical analysis steps. Our initial system properly annotates approximately 88% of standard Korean sentences, and this annotation rate holds across text domains. Previously, the contextual exploration method has been applied successfully to languages as diverse as French and Arabic. Given our success with Korean, we believe that this method can be applied to other agglutinative languages such as Japanese, Turkish and Finnish.