Probabilistic protein function prediction from heterogeneous genome-wide data.

Probabilistic protein function prediction from heterogeneous genome-wide data.
复制标题

根据异质全基因组数据进行概率蛋白质功能预测。

DOI:
10.1371/journal.pone.0000337
复制
发表时间:
2007-03-28
期刊:
影响因子:
3.7
通讯作者:
Kasif, Simon
Kasif, Simon
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Nariai, Naoki;Kolaczyk, Eric D.;Kasif, Simon

文献摘要

参考文献

被引文献

相似文献

高通量测序技术的巨大进步导致了预测基因数量的惊人增长。然而,这些新发现的基因中有很大一部分没有功能分配。幸运的是,各种新颖的高通量全基因组功能筛选技术提供了揭示基因功能的重要线索。整合不同种类的数据来预测蛋白质的功能已经被证明可以提高自动化基因注释系统的准确性。本文提出并评价了一种结合蛋白质-蛋白质相互作用(PPI)数据、基因表达数据、蛋白质基序信息、突变表型数据和蛋白质定位数据的蛋白质功能预测的概率方法。首先,从PPI数据和基因表达数据构建功能连锁图,其中节点(蛋白质)之间的边代表功能相似性的证据。这里的假设是,与不是邻居的蛋白质相比,图形邻居更有可能共享蛋白质的功能。然后,将功能连锁图模型与蛋白质结构域、突变表型和蛋白质定位数据结合使用,以产生功能预测。我们的方法被应用于酿酒酵母基因的功能预测,使用基因本体论(GO)术语作为我们注释的基础。在交叉验证研究中,我们表明,与单独使用PPI数据相比,在50%的准确率下,集成模型的召回率提高了18%。我们还表明,综合预报器明显好于每个单独的预报器。然而,观察到的相对于PPI的改善取决于新的数据来源和要预测的功能类别。令人惊讶的是,在某些情况下,整合会损害整体预测的准确性。最后,我们为目前没有指定功能的463个蛋白质提供了一个假设的GO术语的全面分配。
Dramatic improvements in high throughput sequencing technologies have led to a staggering growth in the number of predicted genes. However, a large fraction of these newly discovered genes do not have a functional assignment. Fortunately, a variety of novel high-throughput genome-wide functional screening technologies provide important clues that shed light on gene function. The integration of heterogeneous data to predict protein function has been shown to improve the accuracy of automated gene annotation systems. In this paper, we propose and evaluate a probabilistic approach for protein function prediction that integrates protein-protein interaction (PPI) data, gene expression data, protein motif information, mutant phenotype data, and protein localization data. First, functional linkage graphs are constructed from PPI data and gene expression data, in which an edge between nodes (proteins) represents evidence for functional similarity. The assumption here is that graph neighbors are more likely to share protein function, compared to proteins that are not neighbors. The functional linkage graph model is then used in concert with protein domain, mutant phenotype and protein localization data to produce a functional prediction. Our method is applied to the functional prediction of Saccharomyces cerevisiae genes, using Gene Ontology (GO) terms as the basis of our annotation. In a cross validation study we show that the integrated model increases recall by 18%, compared to using PPI data alone at the 50% precision. We also show that the integrated predictor is significantly better than each individual predictor. However, the observed improvement vs. PPI depends on both the new source of data and the functional category to be predicted. Surprisingly, in some contexts integration hurts overall prediction accuracy. Lastly, we provide a comprehensive assignment of putative GO terms to 463 proteins that currently have no assigned function.
Pfam:氏族、网络工具和服务。
DOI: 10.1093/nar/gkj149
发表时间: 2006-01-01
影响因子: 14.9
作者:
Finn, Robert D.;Mistry, Jaina;Schuster-Bockler, Benjamin;Griffiths-Jones, Sam;Hollich, Volker;Lassmann, Timo;Moxon, Simon;Marshall, Mhairi;Khanna, Ajay;Durbin, Richard;Eddy, Sean R.;Sonnhammer, Erik L. L.;Bateman, Alex
通讯作者: Bateman, Alex
DOI: 10.1186/gb-2003-4-3-r23
发表时间: 2003
期刊: Genome biology
影响因子: 12.3
作者:
Breitkreutz BJ;Stark C;Tyers M
通讯作者: Tyers M
DOI: 10.1073/pnas.0307326101
发表时间: 2004-03-02
影响因子: 11.1
作者:
Karaoz, U;Murali, TM;Kasif, S
通讯作者: Kasif, S
DOI: 10.1093/bioinformatics/bth294
发表时间: 2004-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Lanckriet, GRG;De Bie, T;Noble, WS
通讯作者: Noble, WS
DOI: 10.1038/47048
发表时间: 1999-11-04
期刊: NATURE
影响因子: 64.8
作者:
Marcotte, EM;Pellegrini, M;Eisenberg, D
通讯作者: Eisenberg, D