How to make the most of NE dictionaries in statistical NER.

How to make the most of NE dictionaries in statistical NER.
复制标题

DOI:
10.1186/1471-2105-9-s11-s5
复制
发表时间:
2008-11-19
期刊:
影响因子:
3
通讯作者:
Ananiadou S
Ananiadou S
中科院分区:
生物学4区
文献类型:
--
作者:
Sasaki Y;Tsuruoka Y;McNaught J;Ananiadou S

文献摘要

被引文献

相似文献

当术语的歧义性和可变性非常高时,即使有大规模的术语资源,基于词典的命名实体识别(NER)也不是一个理想的解决方案。许多关于统计净入学率的研究都试图科普这些问题。然而,如何在统计NER中利用现有的和附加的命名实体(NE)字典并不简单。据推测,将内斯添加到NE字典导致更好的性能。然而,在现实中,需要重新训练NER模型来实现这一目标。我们选择蛋白质名称识别作为案例研究,因为它最遭受的问题与严重的长期变化和歧义。我们已经建立了一种新的方法来提高NER性能,通过添加内斯的NE字典没有再训练。在我们的方法中,第一,已知的内斯被确定在并行的词性(POS)标记的基础上的一般单词字典和NE字典。然后,在POS/PROTEIN标签器输出上训练统计NER,并附上正确的NE标签。我们评估了我们的NER的标准JNLPBA-2004数据集的性能。在将训练数据中出现的蛋白质名称添加到POS标签词典后,测试集上的F分数从73.14提高到73.78,而无需任何模型再训练。在用测试集蛋白质名称丰富标记字典后,性能进一步提高到78.72。我们的方法在蛋白质名称识别方面表现出了很高的性能,这表明如何在统计NER中充分利用已知的内斯。
When term ambiguity and variability are very high, dictionary-based Named Entity Recognition (NER) is not an ideal solution even though large-scale terminological resources are available. Many researches on statistical NER have tried to cope with these problems. However, it is not straightforward how to exploit existing and additional Named Entity (NE) dictionaries in statistical NER. Presumably, addition of NEs to an NE dictionary leads to better performance. However, in reality, the retraining of NER models is required to achieve this. We chose protein name recognition as a case study because it most suffers the problems related to heavy term variation and ambiguity. We have established a novel way to improve the NER performance by adding NEs to an NE dictionary without retraining. In our approach, first, known NEs are identified in parallel with Part-of-Speech (POS) tagging based on a general word dictionary and an NE dictionary. Then, statistical NER is trained on the POS/PROTEIN tagger outputs with correct NE labels attached. We evaluated performance of our NER on the standard JNLPBA-2004 data set. The F-score on the test set has been improved from 73.14 to 73.78 after adding protein names appearing in the training data to the POS tagger dictionary without any model retraining. The performance further increased to 78.72 after enriching the tagging dictionary with test set protein names. Our approach has demonstrated high performance in protein name recognition, which indicates how to make the most of known NEs in statistical NER.