Discovering motif pairs at interaction sites from protein sequences on a proteome-wide scale

Discovering motif pairs at interaction sites from protein sequences on a proteome-wide scale
复制标题

DOI:
10.1093/bioinformatics/btl020
复制
发表时间:
2006-04-15
期刊:
影响因子:
5.8
通讯作者:
Wong, LS
Wong, LS
中科院分区:
生物学3区
文献类型:
--
作者:
Li, HQ;Li, JY;Wong, LS

文献摘要

被引文献

相似文献

动机:蛋白质-蛋白质相互作用由蛋白质相互作用位点介导,是细胞中许多功能过程所固有的。在本文中,我们提出了一种新的方法来发现蛋白质相互作用位点的模式。我们从蛋白质相互作用网络中观察到,存在一种重要的亚结构,称为相互作用蛋白质组对,它表现出这种对中的两个蛋白质集之间的全对全相互作用。这对蛋白质之间的完全相互作用表明了一种共同的相互作用机制,可以称之为相互作用类型。在蛋白质组对的相互作用部位的基序对可以用来表示这种相互作用类型,每个基序通过标准基序发现算法从蛋白质组的序列中得到。从大型蛋白质相互作用网络中系统地发现所有相互作用的蛋白质组是一个具有计算挑战性的问题。通过仔细而复杂的问题转换,使用数据挖掘中广泛研究的频繁模式挖掘的高效算法解决了该问题。结果:我们从酵母相互作用数据集中发现了5349对相互作用的蛋白质组。组内序列同源性的期望值仅为7.48%,表明这些蛋白组内不具有同源性。我们从这些基序对中获得了5343个基序对,以块的形式表示。将我们的基序与块和印刷品数据库中的结构域进行比较,我们发现我们的块可以映射到这两个数据库中平均3.08个相关的块。在这两个数据库中,映射的块出现在总共6794个结构域(蛋白质组)中的4221个。将我们的基序对与来自PDB的3045个相互作用结构域对组成的IPfam进行比较,我们发现在105个不同的PDB复合体中有47个匹配。与另一个可能的领域交互数据库InterDom进行比较,我们发现了203个匹配。
Motivation: Protein-protein interaction, mediated by protein interaction sites, is intrinsic to many functional processes in the cell. In this paper, we propose a novel method to discover patterns in protein interaction sites. We observed from protein interaction networks that there exist a kind of significant substructures called interacting protein group pairs, which exhibit an all-versus-all interaction between the two protein-sets in such a pair. The full-interaction between the pair indicates a common interaction mechanism shared by the proteins in the pair, which can be referred as an interaction type. Motif pairs at the interaction sites of the protein group pairs can be used to represent such interaction type, with each motif derived from the sequences of a protein group by standard motif discovery algorithms. The systematic discovery of all pairs of interacting protein groups from large protein interaction networks is a computationally challenging problem. By a careful and sophisticated problem transformation, the problem is solved using efficient algorithms for mining frequent patterns, a problem extensively studied in data mining.Results: We found 5349 pairs of interacting protein groups from a yeast interaction dataset. The expected value of sequence identity within the groups is only 7.48%, indicating non-homology within these protein groups. We derived 5343 motif pairs from these group pairs, represented in the form of blocks. Comparing our motifs with domains in the BLOCKS and PRINTS databases, we found that our blocks could be mapped to an average of 3.08 correlated blocks in these two databases. The mapped blocks occur 4221 out of total 6794 domains (protein groups) in these two databases. Comparing our motif pairs with iPfam consisting of 3045 interacting domain pairs derived from PDB, we found 47 matches occurring in 105 distinct PDB complexes. Comparing with another putative domain interaction database InterDom, we found 203 matches.