Various criteria in the evaluation of biomedical named entity recognition.

Various criteria in the evaluation of biomedical named entity recognition.
复制标题

DOI:
10.1186/1471-2105-7-92
复制
发表时间:
2006-02-24
期刊:
影响因子:
3
通讯作者:
Hsu WL
Hsu WL
中科院分区:
生物学4区
文献类型:
--
作者:
Tsai RT;Wu SH;Chou WC;Lin YC;He D;Hsiang J;Sung TY;Hsu WL

文献摘要

参考文献

被引文献

相似文献

生物医学领域的文本挖掘正受到越来越多的关注。这一过程的一个关键组成部分是实体识别(NER)。一般来说,两个标注语料库GENIA和GENETAG最常用于训练和测试生物医学命名实体识别(Bio-NER)系统。JNLPBA和BioCreAtivE是使用这些语料库的两个主要的Bio-ner任务。这两个任务对语料库标注采取不同的方法,并使用不同的匹配标准来评估系统性能。本文详细介绍了这些差异,并描述了其他标准。然后,通过对参与上述两个任务的系统进行重新测试,考察了不同的标准和注释方案对系统性能的影响。为了分析JNLPBA和BioCreAtivE的评价的差异,我们进行了实验1,使用BioCreAtivE的分类方案对排名前四的JNLPBA系统进行评价。然后,我们将它们与排名前四的BioCreAtIVE系统进行比较。其中,三个系统同时参与了这两个任务,并且每个系统在JNLPBA上的F-分数都低于BioCreAtIVE。在实验2中,我们使用假设检验和相关系数来寻找BioCreAtIVE的评估方案的替代方案。结果表明,右匹配和左匹配标准与BioCreAtIVE无显著差异。在实验3中,我们提出了一种定制的松弛匹配准则,该准则使用正确的匹配,并将JNLPBA的五个NE类合并为两个,F-Score为81.5%。在实验4中,我们在顶级JNLPBA系统上评估了从宽松到严格的五个匹配标准的范围,并检查了假阴性的百分比。我们的实验给出了当匹配标准放宽时,准确率、召回率和F分数的相对变化。在许多应用中,生物医学NE可以有几个可接受的标签,这些标签可能只是在它们的左边界或右边界不同。然而,大多数语料库只注释了其中的一个。在我们的实验中,我们发现右匹配和左匹配可以作为JNLPBA和BioCreAtIvE的匹配标准的合适的替代。此外,我们的宽松匹配标准表明,用户可以定义他们自己的宽松标准,更符合他们的应用程序需求。
Text mining in the biomedical domain is receiving increasing attention. A key component of this process is named entity recognition (NER). Generally speaking, two annotated corpora, GENIA and GENETAG, are most frequently used for training and testing biomedical named entity recognition (Bio-NER) systems. JNLPBA and BioCreAtIvE are two major Bio-NER tasks using these corpora. Both tasks take different approaches to corpus annotation and use different matching criteria to evaluate system performance. This paper details these differences and describes alternative criteria. We then examine the impact of different criteria and annotation schemes on system performance by retesting systems participated in the above two tasks. To analyze the difference between JNLPBA's and BioCreAtIvE's evaluation, we conduct Experiment 1 to evaluate the top four JNLPBA systems using BioCreAtIvE's classification scheme. We then compare them with the top four BioCreAtIvE systems. Among them, three systems participated in both tasks, and each has an F-score lower on JNLPBA than on BioCreAtIvE. In Experiment 2, we apply hypothesis testing and correlation coefficient to find alternatives to BioCreAtIvE's evaluation scheme. It shows that right-match and left-match criteria have no significant difference with BioCreAtIvE. In Experiment 3, we propose a customized relaxed-match criterion that uses right match and merges JNLPBA's five NE classes into two, which achieves an F-score of 81.5%. In Experiment 4, we evaluate a range of five matching criteria from loose to strict on the top JNLPBA system and examine the percentage of false negatives. Our experiment gives the relative change in precision, recall and F-score as matching criteria are relaxed. In many applications, biomedical NEs could have several acceptable tags, which might just differ in their left or right boundaries. However, most corpora annotate only one of them. In our experiment, we found that right match and left match can be appropriate alternatives to JNLPBA and BioCreAtIvE's matching criteria. In addition, our relaxed-match criterion demonstrates that users can define their own relaxed criteria that correspond more realistically to their application requirements.
DOI: 10.1186/1471-2105-6-s1-s12
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Colosimo ME;Morgan AA;Yeh AS;Colombe JB;Hirschman L
通讯作者: Hirschman L
Genetag:一种名为实体识别的基因/蛋白质的标记语料库。
DOI: 10.1186/1471-2105-6-s1-s3
发表时间: 2005
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Tanabe, L;Xie, N;Thom, LH;Matten, W;Wilbur, WJ
通讯作者: Wilbur, WJ
DOI: 10.1016/j.compbiolchem.2004.09.010
发表时间: 2004-12-01
影响因子: 3.1
作者:
Hu, ZZ;Mani, I;Wu, CH
通讯作者: Wu, CH
使用分类器集合从文本中识别蛋白质/基因名称。
DOI: 10.1186/1471-2105-6-s1-s7
发表时间: 2005
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Zhou, GD;Shen, D;Zhang, J;Su, J;Tan, SH
通讯作者: Tan, SH
DOI: 10.1093/bioinformatics/btg160
发表时间: 2003-07-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Chiang, JH;Yu, HC
通讯作者: Yu, HC