Khmer POS Tagger: A Transformation-based Approach with Hybrid Unknown Word Handling

Khmer POS Tagger: A Transformation-based Approach with Hybrid Unknown Word Handling
复制标题

DOI:
10.1109/icsc.2007.104
复制
发表时间:
2007-09
期刊:
International Conference on Semantic Computing (ICSC 2007)
影响因子:
--
通讯作者:
Chenda Nou;W. Kameyama
Chenda Nou;W. Kameyama
中科院分区:
其他
文献类型:
--
作者:
Chenda Nou;W. Kameyama

文献摘要

被引文献

相似文献

本文对高棉语词性标注进行了初步研究。我们对基于转换的方法的规则算法的应用提出了一些修改,以适应在形态和语法上与英语不同的高棉语。此外,为了克服基于规则的方法在处理未登录词的覆盖范围有限,我们提出了一种混合的方法,结合联合收割机的规则和三元模型。虽然在一个非常小的语料库上进行训练,但这两种方法都比传统方法获得了更高的准确率。该标注器在训练数据和测试数据上分别达到了95.27%和91.96%,其中包含9%的未登录词。
This paper presents an initiative research on Khmer part-of-speech tagger. We propose some modifications on applying rule algorithm of the transformation-based approach to adapt to Khmer language which is morphologically and syntactically different from the English language. Furthermore, to overcome the limited coverage of the rule-based approach in handling unknown words, we propose a hybrid approach to combine the rule-based and trigram models. Although training on a very small corpus, both proposed approaches achieve higher accuracy than the conventional methods. The tagger achieves 95.27% on training data and 91.96% on test data which includes 9% of unknown words.