The development of PIPA: an integrated and automated pipeline for genome-wide protein function annotation.

The development of PIPA: an integrated and automated pipeline for genome-wide protein function annotation.
复制标题

DOI:
10.1186/1471-2105-9-52
复制
发表时间:
2008-01-25
期刊:
影响因子:
3
通讯作者:
Reifman J
Reifman J
中科院分区:
生物学4区
文献类型:
--
作者:
Yu C;Zavaljevski N;Desai V;Johnson S;Stevens FJ;Reifman J

文献摘要

参考文献

被引文献

相似文献

需要自动化蛋白质功能预测方法来跟上高通量测序的步伐。由于存在许多用于推断不同蛋白质功能的程序和数据库,适当整合这些资源的管道将受益于每种方法的优点。然而,集成系统通常不提供生成定制数据库以预测特定蛋白质功能的机制。在这里,我们描述了一个名为PIPA(蛋白质注释管道)的工具,它具有这些功能。PIPA通过将多个程序和数据库(如InterPro和保守结构域数据库)的结果结合到共同的基因本体(GO)术语中来注释蛋白质功能。在PIPA中实现的主要算法是:(1)概况数据库生成算法,其生成定制的概况数据库以预测特定的蛋白质功能,(2)自动本体映射生成算法,其将各种分类方案映射到GO中,以及(3)共识算法,以协调来自集成程序和数据库的注释。采用PIPA的图谱生成算法构建酶图谱数据库CatFam,该数据库预测由酶委员会(EC)编号描述的催化功能。验证测试表明,CatFam的平均召回率和准确率大于95.0%。CatFam与PIPA集成。我们使用关联规则挖掘算法自动生成两个本体之间的映射,从注释的样本蛋白质。将本体的层次拓扑结构嵌入到算法中,增加了生成的映射的数量。特别是,它产生40.0%的额外映射从群集的正交群(COG)EC号码和六倍增加的映射从COG GO条款。对EC数的映射显示出非常高的精确度(99.8%)和召回率(96.6%),而对GO术语的映射显示出中等的精确度(80.0%)和低召回率(33.0%)。我们的共识算法GO注释是基于计算和传播的可能性分数与GO条款。测试结果表明,对于一个给定的召回,共识算法的应用程序产生更高的精度比当共识不使用。PIPA中实现的算法基于来自多个资源的协调预测提供自动化全基因组蛋白质功能注释。
Automated protein function prediction methods are needed to keep pace with high-throughput sequencing. With the existence of many programs and databases for inferring different protein functions, a pipeline that properly integrates these resources will benefit from the advantages of each method. However, integrated systems usually do not provide mechanisms to generate customized databases to predict particular protein functions. Here, we describe a tool termed PIPA (Pipeline for Protein Annotation) that has these capabilities. PIPA annotates protein functions by combining the results of multiple programs and databases, such as InterPro and the Conserved Domains Database, into common Gene Ontology (GO) terms. The major algorithms implemented in PIPA are: (1) a profile database generation algorithm, which generates customized profile databases to predict particular protein functions, (2) an automated ontology mapping generation algorithm, which maps various classification schemes into GO, and (3) a consensus algorithm to reconcile annotations from the integrated programs and databases. PIPA's profile generation algorithm is employed to construct the enzyme profile database CatFam, which predicts catalytic functions described by Enzyme Commission (EC) numbers. Validation tests show that CatFam yields average recall and precision larger than 95.0%. CatFam is integrated with PIPA. We use an association rule mining algorithm to automatically generate mappings between terms of two ontologies from annotated sample proteins. Incorporating the ontologies' hierarchical topology into the algorithm increases the number of generated mappings. In particular, it generates 40.0% additional mappings from the Clusters of Orthologous Groups (COG) to EC numbers and a six-fold increase in mappings from COG to GO terms. The mappings to EC numbers show a very high precision (99.8%) and recall (96.6%), while the mappings to GO terms show moderate precision (80.0%) and low recall (33.0%). Our consensus algorithm for GO annotation is based on the computation and propagation of likelihood scores associated with GO terms. The test results suggest that, for a given recall, the application of the consensus algorithm yields higher precision than when consensus is not used. The algorithms implemented in PIPA provide automated genome-wide protein function annotation based on reconciled predictions from multiple resources.
DOI: 10.1093/nar/gkj095
发表时间: 2006-01-01
影响因子: 14.9
作者:
Maltsev, Natalia;Glass, Elizabeth;Sulakhe, Dinanath;Rodriguez, Alexis;Syed, Mustafa H.;Bompada, Tanuja;Zhang, Yi;D'Souza, Mark
通讯作者: D'Souza, Mark
Pfam:氏族、网络工具和服务。
DOI: 10.1093/nar/gkj149
发表时间: 2006-01-01
影响因子: 14.9
作者:
Finn, Robert D.;Mistry, Jaina;Schuster-Bockler, Benjamin;Griffiths-Jones, Sam;Hollich, Volker;Lassmann, Timo;Moxon, Simon;Marshall, Mhairi;Khanna, Ajay;Durbin, Richard;Eddy, Sean R.;Sonnhammer, Erik L. L.;Bateman, Alex
通讯作者: Bateman, Alex
DOI: 10.1093/nar/gkh044
发表时间: 2004-01-01
影响因子: 14.9
作者:
Hulo, N;Sigrist, CJA;Bairoch, A
通讯作者: Bairoch, A
DOI: 10.1093/bioinformatics/bth021
发表时间: 2004-01-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Sjölander, K
通讯作者: Sjölander, K
DOI: 10.1093/nar/gki034
发表时间: 2005-01-01
影响因子: 14.9
作者:
Bru C;Courcelle E;Carrère S;Beausse Y;Dalmar S;Kahn D
通讯作者: Kahn D