Improving the Scalability of Semi-Markov Conditional Random Fields for Named Entity Recognition

Improving the Scalability of Semi-Markov Conditional Random Fields for Named Entity Recognition
复制标题

DOI:
10.3115/1220175.1220234
复制
发表时间:
2006-07
期刊:
--
影响因子:
--
通讯作者:
Daisuke Okanohara;Yusuke Miyao;Yoshimasa Tsuruoka;Junichi Tsujii
Daisuke Okanohara;Yusuke Miyao;Yoshimasa Tsuruoka;Junichi Tsujii
中科院分区:
其他
文献类型:
--
作者:
Daisuke Okanohara;Yusuke Miyao;Yoshimasa Tsuruoka;Junichi Tsujii

文献摘要

被引文献

相似文献

本文提出了将半正则表达式应用于命名实体识别任务的技术,具有可处理的计算成本。我们的框架可以处理具有长命名实体和许多标签的NER任务,这增加了计算成本。为了降低计算成本,我们提出了两种技术:第一种是使用特征森林,它使我们能够打包特征等效状态;第二种是引入过滤过程,它可以显着减少候选状态的数量。这个框架允许我们使用从基于块的表示中提取的一组丰富的特征,这些特征可以捕获实体的信息特征。我们还介绍了一个简单的技巧,通过将标签信息嵌入到非实体标签中来传递关于远程实体的信息。实验结果表明,在不使用任何外部资源和后处理技术的情况下,我们的模型在JNLPBA 2004共享任务上获得了71.48%的f分。
This paper presents techniques to apply semi-CRFs to Named Entity Recognition tasks with a tractable computational cost. Our framework can handle an NER task that has long named entities and many labels which increase the computational cost. To reduce the computational cost, we propose two techniques: the first is the use of feature forests, which enables us to pack feature-equivalent states, and the second is the introduction of a filtering process which significantly reduces the number of candidate states. This framework allows us to use a rich set of features extracted from the chunk-based representation that can capture informative characteristics of entities. We also introduce a simple trick to transfer information about distant entities by embedding label information into non-entity labels. Experimental results show that our model achieves an F-score of 71.48% on the JNLPBA 2004 shared task without using any external resources or post-processing techniques.