An Error-Driven Word-Character Hybrid Model for Joint Chinese Word Segmentation and POS Tagging

An Error-Driven Word-Character Hybrid Model for Joint Chinese Word Segmentation and POS Tagging
复制标题

DOI:
10.3115/1687878.1687951
复制
发表时间:
2009-08
期刊:
--
影响因子:
--
通讯作者:
Canasai Kruengkrai;Kiyotaka Uchimoto;Jun'ichi Kazama;Yio Wang;Kentaro Torisawa;H. Isahara
Canasai Kruengkrai;Kiyotaka Uchimoto;Jun'ichi Kazama;Yio Wang;Kentaro Torisawa;H. Isahara
中科院分区:
其他
文献类型:
--
作者:
Canasai Kruengkrai;Kiyotaka Uchimoto;Jun'ichi Kazama;Yio Wang;Kentaro Torisawa;H. Isahara

文献摘要

被引文献

相似文献

本文提出了一种汉语分词和词性标注相结合的辨别式词-字混合模型。我们的单词-字符混合模型提供了高性能,因为它既可以处理已知单词,也可以处理未知单词。我们描述了我们的策略,这些策略在学习已知和未知单词的特征方面取得了很好的平衡,并提出了一种错误驱动的策略,该策略通过从训练语料库中的特定错误中获取未知单词的示例来实现这种平衡。我们描述了一种基于保证金注入松弛算法(MIRA)的高效模型训练框架,并在宾夕法尼亚大学中文树库上对我们的方法进行了评估,结果表明,与文献中报道的最新方法相比,我们的方法取得了更好的性能。
In this paper, we present a discriminative word-character hybrid model for joint Chinese word segmentation and POS tagging. Our word-character hybrid model offers high performance since it can handle both known and unknown words. We describe our strategies that yield good balance for learning the characteristics of known and unknown words and propose an error-driven policy that delivers such balance by acquiring examples of unknown words from particular errors in a training corpus. We describe an efficient framework for training our model based on the Margin Infused Relaxed Algorithm (MIRA), evaluate our approach on the Penn Chinese Treebank, and show that it achieves superior performance compared to the state-of-the-art approaches reported in the literature.