Harmonization of gene/protein annotations: towards a gold standard MEDLINE

Harmonization of gene/protein annotations: towards a gold standard MEDLINE
复制标题

DOI:
10.1093/bioinformatics/bts125
复制
发表时间:
2012-05-01
期刊:
影响因子:
5.8
通讯作者:
Rebholz-Schuhmann, Dietrich
Rebholz-Schuhmann, Dietrich
中科院分区:
生物学3区
文献类型:
--
作者:
Campos, David;Matos, Sergio;Rebholz-Schuhmann, Dietrich

文献摘要

被引文献

相似文献

动机:命名实体(NER)的识别是生物医学文本挖掘中的一项基本任务。近年来,利用现有的标注语料库、术语资源和机器学习技术提出了一些NER解决方案。目前,性能最好的解决方案结合了针对单个语料库测量的选定标注解决方案的输出。然而,很少有人对协调标注结果的方法进行系统分析,并与黄金标准语料库(GSCS)的组合进行比较。结果:我们提出了一种机器学习解决方案Totom,它协调了由异质NER解决方案提供的基因/蛋白质标注。它已经针对手动管理的GSC的组合进行了优化和测量。实验表明,我们的方法将最新解决方案的F度量在精确对齐方面提高了10%(接近70%),在嵌套对齐方面提高了22%(接近82%)。我们证明,我们的解决方案提供了跨GSC的可靠标注结果,这是对MEDLINE摘要的同构标注的重要贡献。
Motivation: The recognition of named entities (NER) is an elementary task in biomedical text mining. A number of NER solutions have been proposed in recent years, taking advantage of available annotated corpora, terminological resources and machine- learning techniques. Currently, the best performing solutions combine the outputs from selected annotation solutions measured against a single corpus. However, little effort has been spent on a systematic analysis of methods harmonizing the annotation results and measuring against a combination of Gold Standard Corpora (GSCs).Results: We present Totum, a machine learning solution that harmonizes gene/protein annotations provided by heterogeneous NER solutions. It has been optimized and measured against a combination of manually curated GSCs. The performed experiments show that our approach improves the F-measure of state-of-the-art solutions by up to 10% (achieving approximate to 70%) in exact alignment and 22% (achieving approximate to 82%) in nested alignment. We demonstrate that our solution delivers reliable annotation results across the GSCs and it is an important contribution towards a homogeneous annotation of MEDLINE abstracts.