Extended Models and Tools for High-performance Part-of-speech

Extended Models and Tools for High-performance Part-of-speech
复制标题

用于高性能词性的扩展模型和工具

DOI:
10.3115/990820.990824
复制
发表时间:
2000
期刊:
2007 40th Annual Hawaii International Conference on System Sciences (HICSS'07)
影响因子:
--
通讯作者:
Yuji Matsumoto
Yuji Matsumoto
中科院分区:
--
文献类型:
--
作者:
Masayuki Asahara;Yuji Matsumoto

文献摘要

被引文献

相似文献

统计词性标注在大规模人工标注语料库的基础上具有较高的准确性和鲁棒性。然而,学习模型的增强是必要的,以实现更好的性能。我们正在为日本的形态分析仪ChaSen开发一个学习工具。目前,我们使用的是一个细粒度的POS标签集,大约有500个标签。为了在标签集上应用正常的三元模型,我们需要不切实际的语料库大小。甚至,对于二元语法模型,当我们将所有标签视为不同时,我们无法准备中等大小的注释语料库。科普这种细粒度标记的常用技术是通过将标记集分组到等价类中来减小标记集的大小。我们引入了位置分组的概念,其中标签集在马尔可夫模型中的条件概率中的每个位置处被划分为不同的等价类。此外,为了科普异常现象引起的数据稀疏问题,我们引入了其他几种技术,如词级统计,词级和POS级统计的平滑和选择性三元模型。为了帮助用户确定概率参数,我们引入了一个错误驱动的参数选择方法。然后,我们给出的实验结果,看看应用到现有的日本形态分析仪的工具的效果。
Statistical part-of-speech (POS) taggers achieve high accuracy and robustness when based on large scale manually tagged corpora. However, enhancements of the learning models are necessary to achieve better performance. We are developing a learning tool for a Japanese morphological analyzer called ChaSen. Currently we use a fine-grained POS tag set with about 500 tags. To apply a normal tri gram model on the tag set, we need unrealistic size of corpora. Even, for a bi-gram model, we cannot prepare a moderate size of an annotated corpus, when we take all the tags as distinct. A usual technique to cope with such fine-grained tags is to reduce the size of the tag set by grouping the set of tags into equivalence classes. We introduce the concept of position-wise grouping where the tag set is partitioned into different equivalence classes at each position in the conditional probabilities in the Markov Model. Moreover, to cope with the data sparseness problem caused by exceptional phenomena, we introduce several other techniques such as word-level statistics, smoothing of word-level and POS-level statistics and a selective tri-gram model. To help users determine probabilistic parameters, we introduce an error-driven method for the parameter selection. We then give results of experiments to see the effect of the tools applied to an existing Japanese morphological analyzer.