Improved prediction of protein-protein binding sites using a support vector machines approach

Improved prediction of protein-protein binding sites using a support vector machines approach
复制标题

DOI:
10.1093/bioinformatics/bti242
复制
发表时间:
2005-04-15
期刊:
影响因子:
5.8
通讯作者:
Westhead, DR
Westhead, DR
中科院分区:
生物学3区
文献类型:
--
作者:
Bradford, JR;Westhead, DR

文献摘要

被引文献

相似文献

动机:结构基因组学项目开始产生未知功能的蛋白质结构,因此,如果要在合理的时间内对所有这些结构进行适当的注释,就需要准确的、自动化的蛋白质功能预报器。识别两个相互作用的蛋白质之间的界面可以为蛋白质的功能提供重要线索,并可以减少对接算法预测复合体结构所需的搜索空间。结果:我们将支持向量机(SVM)方法与表面斑块分析相结合来预测蛋白质-蛋白质结合位点。使用留一法交叉验证程序,我们能够成功地预测由具有瞬时和专有界面的蛋白质组成的数据集的76%上的结合位点的位置。在异质交叉验证中,我们训练瞬时复合体上的支持向量机预测专有复合体(反之亦然),我们仍然获得了与留一交叉验证相当的成功率,这表明瞬变界面和专有界面之间有足够的属性共享。
Motivation: Structural genomics projects are beginning to produce protein structures with unknown function, therefore, accurate, automated predictors of protein function are required if all these structures are to be properly annotated in reasonable time. Identifying the interface between two interacting proteins provides important clues to the function of a protein and can reduce the search space required by docking algorithms to predict the structures of complexes.Results: We have combined a support vector machine (SVM) approach with surface patch analysis to predict protein-protein binding sites. Using a leave-one-out cross-validation procedure, we were able to successfully predict the location of the binding site on 76% of our dataset made up of proteins with both transient and obligate interfaces. With heterogeneous cross-validation, where we trained the SVM on transient complexes to predict on obligate complexes (and vice versa), we still achieved comparable success rates to the leave-one-out cross-validation suggesting that sufficient properties are shared between transient and obligate interfaces.