LINNAEUS: a species name identification system for biomedical literature.

LINNAEUS: a species name identification system for biomedical literature.
复制标题

DOI:
10.1186/1471-2105-11-85
复制
发表时间:
2010-02-11
期刊:
影响因子:
3
通讯作者:
Bergman CM
Bergman CM
中科院分区:
生物学4区
文献类型:
--
作者:
Gerner M;Nenadic G;Bergman CM

文献摘要

参考文献

被引文献

相似文献

识别和识别生物医学文献中的物种名称的任务最近被认为是文本和数据挖掘中的一些应用程序,包括基因名称识别,物种特定的文件检索,和语义丰富的生物医学文章的关键。在本文中,我们描述了一个开源的物种名称识别和规范化的软件系统,LINNAEUS,并评估其性能相对于几个自动生成的生物医学语料库,以及一个新的语料库的全文文档手动注释的物种提及。LINNAEUS使用基于字典的方法(作为一个有效的确定性有限状态自动机实现)来识别物种名称和一组语法来解决模糊的提及。与我们的手动注释语料库相比,LINNAEUS在提及级别的召回率为94%,准确率为97%,在文档级别的召回率为98%,准确率为90%。我们的系统成功地解决了不确定物种的歧义问题,PubMed Central全文文档中97%的提及都被解决为明确的NCBI分类标识符。LINNAEUS是一个开源的独立软件系统,能够快速准确地识别和规范物种名称,因此可以集成到一系列生物信息学和文本挖掘应用程序中。该软件和手动注释的语料库可以在http://linnaeus.sourceforge.net/免费下载。
The task of recognizing and identifying species names in biomedical literature has recently been regarded as critical for a number of applications in text and data mining, including gene name recognition, species-specific document retrieval, and semantic enrichment of biomedical articles. In this paper we describe an open-source species name recognition and normalization software system, LINNAEUS, and evaluate its performance relative to several automatically generated biomedical corpora, as well as a novel corpus of full-text documents manually annotated for species mentions. LINNAEUS uses a dictionary-based approach (implemented as an efficient deterministic finite-state automaton) to identify species names and a set of heuristics to resolve ambiguous mentions. When compared against our manually annotated corpus, LINNAEUS performs with 94% recall and 97% precision at the mention level, and 98% recall and 90% precision at the document level. Our system successfully solves the problem of disambiguating uncertain species mentions, with 97% of all mentions in PubMed Central full-text documents resolved to unambiguous NCBI taxonomy identifiers. LINNAEUS is an open source, stand-alone software system capable of recognizing and normalizing species name mentions with speed and accuracy, and can therefore be integrated into a range of bioinformatics and text-mining applications. The software and manually annotated corpus can be downloaded freely at http://linnaeus.sourceforge.net/.
DOI: 10.1093/nar/gkn664
发表时间: 2009-01
影响因子: 14.9
作者:
UniProt Consortium
通讯作者: UniProt Consortium
DOI: 10.1126/science.6189183
发表时间: 1983-01-01
期刊: SCIENCE
影响因子: 56.9
作者:
BARRESINOUSSI, F;CHERMANN, JC;MONTAGNIER, L
通讯作者: MONTAGNIER, L
DOI: 10.1093/bioinformatics/bth386
发表时间: 2004-11-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Corney, DPA;Buxton, BF;Jones, DT
通讯作者: Jones, DT
Oreganno:一种开放访问社区驱动的监管注释资源。
DOI: 10.1093/nar/gkm967
发表时间: 2008-01
影响因子: 14.9
作者:
Griffith OL;Montgomery SB;Bernier B;Chu B;Kasaian K;Aerts S;Mahony S;Sleumer MC;Bilenky M;Haeussler M;Griffith M;Gallo SM;Giardine B;Hooghe B;Van Loo P;Blanco E;Ticoll A;Lithwick S;Portales-Casamar E;Donaldson IJ;Robertson G;Wadelius C;De Bleser P;Vlieghe D;Halfon MS;Wasserman W;Hardison R;Bergman CM;Jones SJ;Open Regulatory Annotation Consortium
通讯作者: Open Regulatory Annotation Consortium
评估生物学的文本挖掘系统:第二次生物综合社区挑战的概述。
DOI: 10.1186/gb-2008-9-s2-s1
发表时间: 2008
期刊: Genome biology
影响因子: 12.3
作者:
Krallinger M;Morgan A;Smith L;Leitner F;Tanabe L;Wilbur J;Hirschman L;Valencia A
通讯作者: Valencia A