NetiNeti: discovery of scientific names from text using machine learning methods.

NetiNeti: discovery of scientific names from text using machine learning methods.
复制标题

DOI:
10.1186/1471-2105-13-211
复制
发表时间:
2012-08-22
期刊:
影响因子:
3
通讯作者:
Miller H
Miller H
中科院分区:
生物学4区
文献类型:
--
作者:
Akella LM;Norton CN;Miller H

文献摘要

参考文献

被引文献

相似文献

一个生物体的学名可以与几乎所有的生物学数据相关联。在许多文本挖掘任务中,名字识别是一个重要的步骤,旨在从生物,生物医学和生物多样性文本源中提取有用的信息。科学名称是连接生物信息的重要元数据元素。我们提出了NetiNeti(从文本信息中提取名称-用于分类索引的名称提取),这是一种基于机器学习的方法,用于识别科学名称,包括从文本中发现新的物种名称,还将处理拼写错误,OCR错误和名称的其他变化。该系统使用科学名称规则生成候选名称,并应用概率机器学习方法根据候选名称的结构特征和从其上下文中派生的特征对名称进行分类。NetiNeti还可以使用上下文信息消除科学名称与其他名称的歧义。我们评估了NetiNeti的遗产生物多样性文本和生物医学文献(MEDLINE)。NetiNeti的表现更好(精度= 98.9%和召回= 70.5%)相比,流行的基于字典的方法(精度= 97.5%和召回= 54.3%)在600页的生物多样性的书,手动标记的注释。在PubMed Central的一小部分标注了科学名称的全文文章上,准确率和召回率分别为98.5%和96.2%。当在完整的MEDLINE数据库上使用时,NetiNeti在超过1,880,000个PubMed记录中发现了超过190,000个独特的二项式和三项式名称。NetiNeti还成功地识别了网页中提到的几乎所有新物种名称。我们提出了NetiNeti,一种基于机器学习的方法,用于识别和发现科学名称。实现该方法的系统可以在http://namefinding.ubio.org上访问。
A scientific name for an organism can be associated with almost all biological data. Name identification is an important step in many text mining tasks aiming to extract useful information from biological, biomedical and biodiversity text sources. A scientific name acts as an important metadata element to link biological information. We present NetiNeti (Name Extraction from Textual Information-Name Extraction for Taxonomic Indexing), a machine learning based approach for recognition of scientific names including the discovery of new species names from text that will also handle misspellings, OCR errors and other variations in names. The system generates candidate names using rules for scientific names and applies probabilistic machine learning methods to classify names based on structural features of candidate names and features derived from their contexts. NetiNeti can also disambiguate scientific names from other names using the contextual information. We evaluated NetiNeti on legacy biodiversity texts and biomedical literature (MEDLINE). NetiNeti performs better (precision = 98.9% and recall = 70.5%) compared to a popular dictionary based approach (precision = 97.5% and recall = 54.3%) on a 600-page biodiversity book that was manually marked by an annotator. On a small set of PubMed Central’s full text articles annotated with scientific names, the precision and recall values are 98.5% and 96.2% respectively. NetiNeti found more than 190,000 unique binomial and trinomial names in more than 1,880,000 PubMed records when used on the full MEDLINE database. NetiNeti also successfully identifies almost all of the new species names mentioned within web pages. We present NetiNeti, a machine learning based approach for identification and discovery of scientific names. The system implementing the approach can be accessed at http://namefinding.ubio.org.
DOI: 10.1023/a:1007506220214
发表时间: 1999-02-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Beeferman, D;Berger, A;Lafferty, J
通讯作者: Lafferty, J
DOI: 10.1093/bioinformatics/btm557
发表时间: 2008-01-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Rebholz-Schuhmann, Dietrich;Arregui, Miguel;Jimeno, Antonio
通讯作者: Jimeno, Antonio
DOI: 10.1093/bioinformatics/btl408
发表时间: 2006-10-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Plake, Conrad;Schiemann, Torsten;Leser, Ulf
通讯作者: Leser, Ulf
DOI: 10.1109/34.588021
发表时间: 1997-04-01
影响因子: 23.6
作者:
DellaPietra, S;DellaPietra, V;Lafferty, J
通讯作者: Lafferty, J
DOI: 10.1093/bioinformatics/btm109
发表时间: 2007-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Leary, Patrick R.;Remsen, David P.;Sarkar, Indra Neil
通讯作者: Sarkar, Indra Neil