Information assessment on predicting protein-protein interactions.

Information assessment on predicting protein-protein interactions.
复制标题

DOI:
10.1186/1471-2105-5-154
复制
发表时间:
2004-10-18
期刊:
影响因子:
3
通讯作者:
Zhao H
Zhao H
中科院分区:
生物学4区
文献类型:
--
作者:
Lin N;Wu B;Jansen R;Gerstein M;Zhao H

文献摘要

参考文献

被引文献

相似文献

识别蛋白质-蛋白质相互作用是理解细胞分子机制的基础。蛋白质-蛋白质相互作用的蛋白质组范围的研究具有重要价值,但是高通量实验技术遭受高的假阳性和假阴性预测率。除了高通量实验数据,许多不同类型的基因组数据可以帮助预测蛋白质-蛋白质相互作用,如mRNA表达,定位,必要性和功能注释。对不同证据的信息贡献进行评估有助于建立具有可比或更好预测精度的更简约的模型,并获得蛋白质-蛋白质相互作用与其他基因组信息之间关系的生物学见解。我们的评估是基于贝叶斯网络方法中使用的基因组特征来预测酵母中全基因组的蛋白质-蛋白质相互作用。在特殊情况下,当一个人没有任何丢失的信息的任何功能,我们的分析表明,有一个更大的信息贡献的功能分类比表达相关性或本质。我们还表明,在这种情况下,替代模型,如逻辑回归和随机森林,可能比贝叶斯网络更有效地预测相互作用。在完全信息子集带来的限制问题中,我们发现MIPS和基因本体论(GO)功能相似性数据集是Jansen等人提出的框架下预测蛋白质-蛋白质相互作用的主要信息贡献者。基于MIPS和GO的随机森林信息单独可以给出高度准确的分类。在这个完整信息的特定子集中,添加其他基因组数据对改善预测几乎没有帮助。我们还发现,贝叶斯方法中使用的数据离散化降低了分类性能。
Identifying protein-protein interactions is fundamental for understanding the molecular machinery of the cell. Proteome-wide studies of protein-protein interactions are of significant value, but the high-throughput experimental technologies suffer from high rates of both false positive and false negative predictions. In addition to high-throughput experimental data, many diverse types of genomic data can help predict protein-protein interactions, such as mRNA expression, localization, essentiality, and functional annotation. Evaluations of the information contributions from different evidences help to establish more parsimonious models with comparable or better prediction accuracy, and to obtain biological insights of the relationships between protein-protein interactions and other genomic information. Our assessment is based on the genomic features used in a Bayesian network approach to predict protein-protein interactions genome-wide in yeast. In the special case, when one does not have any missing information about any of the features, our analysis shows that there is a larger information contribution from the functional-classification than from expression correlations or essentiality. We also show that in this case alternative models, such as logistic regression and random forest, may be more effective than Bayesian networks for predicting interactions. In the restricted problem posed by the complete-information subset, we identified that the MIPS and Gene Ontology (GO) functional similarity datasets as the dominating information contributors for predicting the protein-protein interactions under the framework proposed by Jansen et al. Random forests based on the MIPS and GO information alone can give highly accurate classifications. In this particular subset of complete information, adding other genomic data does little for improving predictions. We also found that the data discretizations used in the Bayesian methods decreased classification performance.
DOI: 10.1038/415141a
发表时间: 2002-01-10
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Bösche, M;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1016/s1097-2765(00)80114-8
发表时间: 1998-07-01
期刊: MOLECULAR CELL
影响因子: 16
作者:
Cho, RJ;Campbell, MJ;Davis, RW
通讯作者: Davis, RW
DOI: 10.1093/nar/30.1.31
发表时间: 2002-01-01
影响因子: 14.9
作者:
Mewes, HW;Frishman, D;Weil, B
通讯作者: Weil, B
DOI: 10.1101/gad.970902
发表时间: 2002-03-15
影响因子: 10.5
作者:
Kumar, A;Agarwal, S;Snyder, M
通讯作者: Snyder, M
DOI: 10.1073/pnas.96.8.4285
发表时间: 1999-04-13
影响因子: 11.1
作者:
Pellegrini, M;Marcotte, EM;Yeates, TO
通讯作者: Yeates, TO