Evaluating gold standard corpora against gene/protein tagging solutions and lexical resources.

Evaluating gold standard corpora against gene/protein tagging solutions and lexical resources.
复制标题

DOI:
10.1186/2041-1480-4-28
复制
发表时间:
2013-10-11
影响因子:
1.9
通讯作者:
Lewin I
Lewin I
中科院分区:
工程技术4区
文献类型:
--
作者:
Rebholz-Schuhmann D;Kafkas S;Kim JH;Li C;Jimeno Yepes A;Hoehndorf R;Backofen R;Lewin I

文献摘要

被引文献

相似文献

从科学文献中识别蛋白质和基因名称(PGN)需要语义资源:术语和词汇资源将候选术语传递到PGN标记解决方案中,金标准语料库(GSC)训练它们识别术语参数和上下文特征。理想情况下,这三种资源(即语料库、词典和标记器)涵盖相同的领域知识,从而支持识别相同类型的pgn并涵盖所有这些知识。不幸的是,这三种资源都不能作为主导标准,因此值得探讨这三种资源如何相互遵从。我们系统地比较了不同的PGN标记器与公开可用的语料库,并分析了包含的词汇资源对其性能的影响。特别是,我们通过假正滤波确定性能增益,这有助于消除识别的pgn的歧义。一般来说,机器学习方法(ML-Tag)用于PGN标记对biocrea2和Jnlpba GSCs(精确匹配)表现出更高的F1-measure性能,而基于词典的方法(LexTag)结合消歧方法在FsuPrge和PennBio上表现出更好的结果。ML-Tag解决方案平衡了精度和召回率,而LexTag解决方案在所有语料库的相同f1测量下具有不同的精度和召回率。更大的词汇资源可以实现更高的召回率,这也会引入更多的噪声(假阳性结果)。如果测试语料库与训练语料库来自相同的GSC,那么ML-Tag解决方案当然表现最好。正如预期的那样,假阴性错误表征了测试语料库,而另一方面,假阳性错误的轮廓表征了标注解决方案。基于大型术语资源并结合假阳性过滤的Lex-Tag解决方案产生更好的结果,此外,与ML-Tag解决方案相比,它还提供了来自知识来源的概念标识符。标准的ml标签解决方案实现了高性能,但不能跨越所有的语料库,因此应该使用几个不同的语料库进行训练,以减少可能的偏差。LexTag解决方案在精度和召回性能方面有不同的配置文件,但具有相似的f1测量。这一结果令人惊讶,表明它们覆盖了最常见的命名标准的一部分,但在语料库中处理术语可变性的方式不同。应用于LexTag解决方案的假阳性过滤确实通过在不显著影响召回率的情况下提高精度来改善结果。标注方案的协调与标注解决方案中的标准化词汇资源相结合,将使它们具有可比性,并为共享标准铺平道路。
The identification of protein and gene names (PGNs) from the scientific literature requires semantic resources: Terminological and lexical resources deliver the term candidates into PGN tagging solutions and the gold standard corpora (GSC) train them to identify term parameters and contextual features. Ideally all three resources, i.e. corpora, lexica and taggers, cover the same domain knowledge, and thus support identification of the same types of PGNs and cover all of them. Unfortunately, none of the three serves as a predominant standard and for this reason it is worth exploring, how these three resources comply with each other. We systematically compare different PGN taggers against publicly available corpora and analyze the impact of the included lexical resource in their performance. In particular, we determine the performance gains through false positive filtering, which contributes to the disambiguation of identified PGNs. In general, machine learning approaches (ML-Tag) for PGN tagging show higher F1-measure performance against the BioCreative-II and Jnlpba GSCs (exact matching), whereas the lexicon based approaches (LexTag) in combination with disambiguation methods show better results on FsuPrge and PennBio. The ML-Tag solutions balance precision and recall, whereas the LexTag solutions have different precision and recall profiles at the same F1-measure across all corpora. Higher recall is achieved with larger lexical resources, which also introduce more noise (false positive results). The ML-Tag solutions certainly perform best, if the test corpus is from the same GSC as the training corpus. As expected, the false negative errors characterize the test corpora and – on the other hand – the profiles of the false positive mistakes characterize the tagging solutions. Lex-Tag solutions that are based on a large terminological resource in combination with false positive filtering produce better results, which, in addition, provide concept identifiers from a knowledge source in contrast to ML-Tag solutions. The standard ML-Tag solutions achieve high performance, but not across all corpora, and thus should be trained using several different corpora to reduce possible biases. The LexTag solutions have different profiles for their precision and recall performance, but with similar F1-measure. This result is surprising and suggests that they cover a portion of the most common naming standards, but cope differently with the term variability across the corpora. The false positive filtering applied to LexTag solutions does improve the results by increasing their precision without compromising significantly their recall. The harmonisation of the annotation schemes in combination with standardized lexical resources in the tagging solutions will enable their comparability and will pave the way for a shared standard.