A Stacked Sub-Word Model for Joint Chinese Word Segmentation and Part-of-Speech Tagging

A Stacked Sub-Word Model for Joint Chinese Word Segmentation and Part-of-Speech Tagging
复制标题

DOI:
--
复制
发表时间:
2011-06
期刊:
--
影响因子:
--
通讯作者:
Weiwei SUN
Weiwei SUN
中科院分区:
其他
文献类型:
--
作者:
Weiwei SUN

文献摘要

被引文献

相似文献

联合分词和词性标注的组合搜索空间大,使得高效解码非常困难。因此,表示丰富上下文的有效高阶特征不方便使用。在这项工作中,我们提出了一种新的堆叠子字模型,这一任务,同时考虑到效率和有效性。我们的解决方案是一个两步的过程。首先,一个基于词的分割器,一个基于字符的分割器和一个本地字符分类器进行训练,以产生粗分割和POS信息。其次,三个预测器的输出被合并成子词序列,这些子词序列被进一步括起来,并由细粒度子词标注器用POS标签进行标注。粗到细的搜索方案是有效的,而在子词标记步骤中,可以近似地导出丰富的上下文特征。在宾夕法尼亚大学中国树库上的评价表明,我们的模型比文献中报道的最好的系统有所改进。
The large combined search space of joint word segmentation and Part-of-Speech (POS) tagging makes efficient decoding very hard. As a result, effective high order features representing rich contexts are inconvenient to use. In this work, we propose a novel stacked subword model for this task, concerning both efficiency and effectiveness. Our solution is a two step process. First, one word-based segmenter, one character-based segmenter and one local character classifier are trained to produce coarse segmentation and POS information. Second, the outputs of the three predictors are merged into sub-word sequences, which are further bracketed and labeled with POS tags by a fine-grained sub-word tagger. The coarse-to-fine search scheme is efficient, while in the sub-word tagging step rich contextual features can be approximately derived. Evaluation on the Penn Chinese Tree-bank shows that our model yields improvements over the best system reported in the literature.