An integrated approach to the prediction of domain-domain interactions.

An integrated approach to the prediction of domain-domain interactions.
复制标题

DOI:
10.1186/1471-2105-7-269
复制
发表时间:
2006-05-25
期刊:
影响因子:
3
通讯作者:
Chen, Ting
Chen, Ting
中科院分区:
生物学4区
文献类型:
--
作者:
Lee, Hyunju;Deng, Minghua;Sun, Fengzhu;Chen, Ting

文献摘要

参考文献

被引文献

相似文献

高通量技术的发展已经产生了多个物种的大规模蛋白质相互作用数据集,并且已经做出了显著的努力来分析数据集以了解蛋白质活性。考虑到蛋白质相互作用的基本单位是结构域相互作用,在结构域水平上理解蛋白质相互作用至关重要。许多不同的生物数据集的可用性提供了一个机会,发现潜在的结构域相互作用蛋白质相互作用,通过这些生物数据集的整合。我们结合联合收割机蛋白质相互作用数据集,从多个物种,分子序列,和基因本体构建一套高置信度域域相互作用。首先,我们提出了一个新的措施,预期数量的相互作用,为每对域,评分域的相互作用的基础上,在一个物种中的蛋白质相互作用的数据,并表明它具有类似的性能,由Riley等人定义的E值。我们的新措施适用于蛋白质相互作用的数据集从酵母,蠕虫,果蝇和人类。其次,在已知的蛋白质中共存的结构域对和具有相同的基因本体论功能注释的结构域对的信息被纳入使用贝叶斯方法构建一个高置信度的域-域相互作用集。最后,我们通过比较预测的结构域相互作用与iPfam数据库中定义的基于蛋白质结构的相互作用来评估结构域之间的相互作用。通过与H.幽门。结果,总共获得了2,391个高置信度的结构域相互作用,这些结构域相互作用被用于解开几种蛋白质复合物中详细的蛋白质和结构域相互作用。我们的研究表明,基于贝叶斯方法的多个生物数据集的整合提供了一个可靠的框架来预测域相互作用。通过集成多个数据源,可以显著提高预测域交互的覆盖率和准确性。
The development of high-throughput technologies has produced several large scale protein interaction data sets for multiple species, and significant efforts have been made to analyze the data sets in order to understand protein activities. Considering that the basic units of protein interactions are domain interactions, it is crucial to understand protein interactions at the level of the domains. The availability of many diverse biological data sets provides an opportunity to discover the underlying domain interactions within protein interactions through an integration of these biological data sets. We combine protein interaction data sets from multiple species, molecular sequences, and gene ontology to construct a set of high-confidence domain-domain interactions. First, we propose a new measure, the expected number of interactions for each pair of domains, to score domain interactions based on protein interaction data in one species and show that it has similar performance as the E-value defined by Riley et al.. Our new measure is applied to the protein interaction data sets from yeast, worm, fruitfly and humans. Second, information on pairs of domains that coexist in known proteins and on pairs of domains with the same gene ontology function annotations are incorporated to construct a high-confidence set of domain-domain interactions using a Bayesian approach. Finally, we evaluate the set of domain-domain interactions by comparing predicted domain interactions with those defined in iPfam database that were derived based on protein structures. The accuracy of predicted domain interactions are also confirmed by comparing with experimentally obtained domain interactions from H. pylori . As a result, a total of 2,391 high-confidence domain interactions are obtained and these domain interactions are used to unravel detailed protein and domain interactions in several protein complexes. Our study shows that integration of multiple biological data sets based on the Bayesian approach provides a reliable framework to predict domain interactions. By integrating multiple data sources, the coverage and accuracy of predicted domain interactions can be significantly increased.
DOI: 10.1093/bioinformatics/btg118
发表时间: 2003-05-22
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Ng, SK;Zhang, Z;Tan, SH
通讯作者: Tan, SH
DOI: 10.1101/gr.153002
发表时间: 2002-10-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Deng, MH;Mehta, S;Chen, T
通讯作者: Chen, T
DOI: 10.1093/nar/30.1.31
发表时间: 2002-01-01
影响因子: 14.9
作者:
Mewes, HW;Frishman, D;Weil, B
通讯作者: Weil, B
DOI: 10.1016/s0968-0004(98)01253-5
发表时间: 1998-09-01
影响因子: 13.8
作者:
Henrick, K;Thornton, JM
通讯作者: Thornton, JM
DOI: 10.1126/science.1091403
发表时间: 2004-01-23
期刊: SCIENCE
影响因子: 56.9
作者:
Li, SM;Armstrong, CM;Vidal, M
通讯作者: Vidal, M