Bidirectional Inference with the Easiest-First Strategy for Tagging Sequence Data

Bidirectional Inference with the Easiest-First Strategy for Tagging Sequence Data
复制标题

DOI:
10.3115/1220575.1220634
复制
发表时间:
2005-10
期刊:
--
影响因子:
--
通讯作者:
Yoshimasa Tsuruoka;Junichi Tsujii
Yoshimasa Tsuruoka;Junichi Tsujii
中科院分区:
其他
文献类型:
--
作者:
Yoshimasa Tsuruoka;Junichi Tsujii

文献摘要

被引文献

相似文献

针对词性标注、命名实体识别、文本组块等序列标注问题,提出了一种双向推理算法。该算法可以枚举所有可能的分解结构,并在多项式时间内找到概率最高的序列及其对应的分解结构。我们还提出了一种基于最简单优先策略的高效译码算法,与全双向推理相比,该算法具有相对较好的性能,且计算量显著降低。词性标注和文本组块的实验结果表明,所提出的双向推理方法的性能一致优于单向推理方法,而双向MEMM的性能与包括核支持向量机在内的最新学习算法的性能相当。
This paper presents a bidirectional inference algorithm for sequence labeling problems such as part-of-speech tagging, named entity recognition and text chunking. The algorithm can enumerate all possible decomposition structures and find the highest probability sequence together with the corresponding decomposition structure in polynomial time. We also present an efficient decoding algorithm based on the easiest-first strategy, which gives comparably good performance to full bidirectional inference with significantly lower computational cost. Experimental results of part-of-speech tagging and text chunking show that the proposed bidirectional inference methods consistently outperform unidirectional inference methods and bidirectional MEMMs give comparable performance to that achieved by state-of-the-art learning algorithms including kernel support vector machines.