Exploiting sequence and structure homologs to identify protein-protein binding sites

Exploiting sequence and structure homologs to identify protein-protein binding sites
复制标题

DOI:
10.1002/prot.20741
复制
发表时间:
2006-02-15
影响因子:
2.9
通讯作者:
Bourne, PE
Bourne, PE
中科院分区:
生物学4区
文献类型:
--
作者:
Chung, JL;Wang, W;Bourne, PE

文献摘要

被引文献

相似文献

实验衍生的三维结构数量的快速增加为更好地理解和随后预测蛋白质-蛋白质相互作用提供了机会。在本研究中,结构保守的残基源自已知复合物的各个组分的多重结构比对,并且根据晶体学 B 因子对分配的保守分数进行加权,以考虑将导致不良比对的结构灵活性。然后将序列图谱和可获取的表面积信息与保守评分相结合,使用支持向量机 (SVM) 预测蛋白质-蛋白质结合位点。保护分数的结合显着提高了支持向量机的性能。大约52%的结合位点被精确预测(该位点中超过70%的残基被识别); 77%的结合位点被正确预测(该位点中超过50%的残基被识别),21%的结合位点被预测的残基部分覆盖(一些残基被识别)。结果支持这样的假设:在许多情况下,蛋白质界面需要一些残基来提供刚性,以最小化复杂形成时的熵成本。
A rapid increase in the number of experimentally derived three-dimensional structures provides an opportunity to better understand and subsequently predict protein-protein interactions. In this study, structurally conserved residues were derived from multiple structure alignments of the individual components of known complexes and the assigned conservation score was weighted based on the crystallographic B factor to account for the structural flexibility that will result in a poor alignment. Sequence profile and accessible surface area information was then combined with the conservation score to predict protein-protein binding sites using a Support Vector Machine (SVM). The incorporation of the conservation score significantly improved the performance of the SVM. About 52% of the binding sites were precisely predicted (greater than 70% of the residues in the site were identified); 77% of the binding sites were correctly predicted (greater than 50% of the residues in the site were identified), and 21% of the binding sites were partially covered by the predicted residues (some residues were identified). The results support the hypothesis that in many cases protein interfaces require some residues to provide rigidity to minimize the entropic cost upon complex formation.