Generalisation in named entity recognition: A quantitative analysis

Generalisation in named entity recognition: A quantitative analysis
复制标题

DOI:
10.1016/j.csl.2017.01.012
复制
发表时间:
2017-07-01
影响因子:
4.3
通讯作者:
Bontcheva, Kalina
Bontcheva, Kalina
中科院分区:
计算机科学3区
文献类型:
--
作者:
Augenstein, Isabelle;Derczynski, Leon;Bontcheva, Kalina

文献摘要

被引文献

相似文献

命名实体识别(NER)是一项关键的NLP任务,在Web和用户生成的内容中,由于其语言的多样性和不断变化,因此更具挑战性。本文旨在量化这种多样性如何影响最先进的NER方法,通过测量命名实体(NE)和上下文的变化,特征稀疏性,以及它们对精度和召回的影响。特别是,我们的研究结果表明,NER方法很难在训练数据有限的情况下概括不同的体裁。特别是,看不见的内斯扮演着重要的角色,在社交媒体等不同类型中的发生率高于新闻专线等更常规的类型。再加上更普遍的不可见特征的发生率更高以及缺乏大型训练语料库,这导致与更常规的类型相比,不同类型的Fl分数显著更低。我们还发现,领先的系统在很大程度上依赖于训练数据中发现的表面形式,在这些之外存在泛化问题,并为这一观察提供了解释。(C)2017作者由Elsevier Ltd.发布。这是CC BY许可下的开放获取文章。
Named Entity Recognition (NER) is a key NLP task, which is all the more challenging on Web and user-generated content with their diverse and continuously changing language. This paper aims to quantify how this diversity impacts state-of-the-art NER methods, by measuring named entity (NE) and context variability, feature sparsity, and their effects on precision and recall. In particular, our findings indicate that NER approaches struggle to generalise in diverse genres with limited training data. Unseen NEs, in particular, play an important role, which have a higher incidence in diverse genres such as social media than in more regular genres such as newswire. Coupled with a higher incidence of unseen features more generally and the lack of large training corpora, this leads to significantly lower Fl scores for diverse genres as compared to more regular ones. We also find that leading systems rely heavily on surface forms found in training data, having problems generalising beyond these, and offer explanations for this observation. (C) 2017 The Authors. Published by Elsevier Ltd. This is an open access article article under the CC BY license.