Integrated protein function prediction by mining function associations, sequences, and protein-protein and gene-gene interaction networks.

Integrated protein function prediction by mining function associations, sequences, and protein-protein and gene-gene interaction networks.
复制标题

DOI:
10.1016/j.ymeth.2015.09.011
复制
发表时间:
2016-01-15
期刊:
Methods (San Diego, Calif.)
影响因子:
--
通讯作者:
Cheng J
Cheng J
中科院分区:
其他
文献类型:
--
作者:
Cao R;Cheng J

文献摘要

被引文献

相似文献

蛋白质功能预测是生物信息学和计算生物学中的一个重要而又具有挑战性的问题。功能相关的生物信息,如蛋白质序列,基因表达和蛋白质-蛋白质相互作用已被大多单独用于蛋白质功能预测。主要挑战之一是如何有效地整合传统信息和新信息的多个来源,例如从染色体构象数据生成的空间基因-基因相互作用网络,以改善蛋白质功能预测。在这项工作中,我们开发了三种不同的概率得分(MIS,SEQ和NET得分)结合联合收割机蛋白质序列,功能协会,蛋白质-蛋白质相互作用和空间基因-基因相互作用网络蛋白质功能预测。MIS评分主要由PSI-BLAST搜索发现的同源蛋白质以及通过挖掘Swiss-Prot数据库学习的Gene Ontology术语之间的关联规则生成。SEQ评分由蛋白质序列生成。NET得分是从蛋白质-蛋白质相互作用和空间基因-基因相互作用网络中产生的。这三个分数结合在一个新的统计多积分评分系统(SMISS)预测蛋白质功能。我们在2011年CAFA(Critical Assessment of Function Annotation)的数据集上测试了SMISS。该方法的性能大大优于三个基线方法和先进的方法的基础上蛋白质的轮廓序列比较,轮廓轮廓比较,和域共现网络根据最大F-措施。
Protein function prediction is an important and challenging problem in bioinformatics and computational biology. Functionally relevant biological information such as protein sequences, gene expression, and protein–protein interactions has been used mostly separately for protein function prediction. One of the major challenges is how to effectively integrate multiple sources of both traditional and new information such as spatial gene–gene interaction networks generated from chromosomal conformation data together to improve protein function prediction. In this work, we developed three different probabilistic scores (MIS, SEQ, and NET score) to combine protein sequence, function associations, and protein–protein interaction and spatial gene–gene interaction networks for protein function prediction. The MIS score is mainly generated from homologous proteins found by PSI-BLAST search, and also association rules between Gene Ontology terms, which are learned by mining the Swiss-Prot database. The SEQ score is generated from protein sequences. The NET score is generated from protein–protein interaction and spatial gene–gene interaction networks. These three scores were combined in a new Statistical Multiple Integrative Scoring System (SMISS) to predict protein function. We tested SMISS on the data set of 2011 Critical Assessment of Function Annotation (CAFA). The method performed substantially better than three base-line methods and an advanced method based on protein profile–sequence comparison, profile–profile comparison, and domain co-occurrence networks according to the maximum F-measure.