How to link ontologies and protein-protein interactions to literature: text-mining approaches and the BioCreative experience.

How to link ontologies and protein-protein interactions to literature: text-mining approaches and the BioCreative experience.
复制标题

DOI:
10.1093/database/bas017
复制
发表时间:
2012
期刊:
Database : the journal of biological databases and curation
影响因子:
--
通讯作者:
Chatr-aryamontri A
Chatr-aryamontri A
中科院分区:
其他
文献类型:
--
作者:
Krallinger M;Leitner F;Vazquez M;Salgado D;Marcelle C;Tyers M;Valencia A;Chatr-aryamontri A

文献摘要

参考文献

被引文献

相似文献

开发本体和受控词汇表以提高手动文献策展的效率和一致性,以实现更正式的生物定位工作流程结果并最终改善生物数据的分析,这一点越来越受到关注。已成功用于此目的的两个本体是用于注释基因产物方面的基因本体(GO)和用于存档蛋白质-蛋白质相互作用的数据库的分子相互作用本体(PSI-MI)。蛋白质相互作用的检查已被证明是非常有前途的细胞过程的理解。从生物医学文献到生物本体术语的信息的手动映射是策展管道中最具挑战性的组成部分之一。它要求专业策展人解释文章中包含的自然语言描述,并在本体(受控词汇表)中推断其语义等价物。由于手动策展是一个耗时的过程,因此有强烈的动机来实现文本挖掘技术,以自动从自由文本中提取注释。已经设计了一系列文本挖掘策略来帮助自动提取生物数据。这些策略或者识别在文献中反复使用的技术术语,并将它们作为包含在本体中的候选项,或者检索用作注释本体术语的证据支持的段落,例如从PSI-MI或GO控制的词汇表。在这里,我们提供了一个总体概述,目前的文本挖掘方法,自动提取注释的GO和PSI-MI本体论术语的背景下,BioCreative(生物学信息提取系统的关键评估)的挑战。特别强调的是蛋白质-蛋白质相互作用的数据和PSI-MI术语指的是相互作用的检测方法。
There is an increasing interest in developing ontologies and controlled vocabularies to improve the efficiency and consistency of manual literature curation, to enable more formal biocuration workflow results and ultimately to improve analysis of biological data. Two ontologies that have been successfully used for this purpose are the Gene Ontology (GO) for annotating aspects of gene products and the Molecular Interaction ontology (PSI-MI) used by databases that archive protein–protein interactions. The examination of protein interactions has proven to be extremely promising for the understanding of cellular processes. Manual mapping of information from the biomedical literature to bio-ontology terms is one of the most challenging components in the curation pipeline. It requires that expert curators interpret the natural language descriptions contained in articles and infer their semantic equivalents in the ontology (controlled vocabulary). Since manual curation is a time-consuming process, there is strong motivation to implement text-mining techniques to automatically extract annotations from free text. A range of text mining strategies has been devised to assist in the automated extraction of biological data. These strategies either recognize technical terms used recurrently in the literature and propose them as candidates for inclusion in ontologies, or retrieve passages that serve as evidential support for annotating an ontology term, e.g. from the PSI-MI or GO controlled vocabularies. Here, we provide a general overview of current text-mining methods to automatically extract annotations of GO and PSI-MI ontology terms in the context of the BioCreative (Critical Assessment of Information Extraction Systems in Biology) challenge. Special emphasis is given to protein–protein interaction data and PSI-MI terms referring to interaction detection methods.
DOI: 10.1186/1471-2105-12-s8-s8
发表时间: 2011-10-03
期刊: BMC bioinformatics
影响因子: 3
作者:
Chatr-Aryamontri A;Winter A;Perfetto L;Briganti L;Licata L;Iannuccelli M;Castagnoli L;Cesareni G;Tyers M
通讯作者: Tyers M
Goannotator:将蛋白质GO注释与证据文本联系起来。
DOI: 10.1186/1747-5333-1-19
发表时间: 2006-12-20
期刊: Journal of biomedical discovery and collaboration
影响因子: --
作者:
Couto, Francisco M;Silva, Mario J;Lee, Vivian;Dimmer, Emily;Camon, Evelyn;Apweiler, Rolf;Kirsch, Harald;Rebholz-Schuhmann, Dietrich
通讯作者: Rebholz-Schuhmann, Dietrich
生物公约概述:生物学信息提取的批判性评估。
DOI: 10.1186/1471-2105-6-s1-s1
发表时间: 2005
期刊: BMC bioinformatics
影响因子: 3
作者:
Hirschman L;Yeh A;Blaschke C;Valencia A
通讯作者: Valencia A
DOI: 10.1093/nar/gkp983
发表时间: 2010-01
影响因子: 14.9
作者:
Ceol A;Chatr Aryamontri A;Licata L;Peluso D;Briganti L;Perfetto L;Castagnoli L;Cesareni G
通讯作者: Cesareni G
DOI: 10.1093/nar/gkq1243
发表时间: 2011-01
影响因子: 14.9
作者:
Galperin MY;Cochrane GR
通讯作者: Cochrane GR