Prediction of protein binding sites in protein structures using hidden Markov support vector machine.

Prediction of protein binding sites in protein structures using hidden Markov support vector machine.
复制标题

使用隐马尔可夫支持向量机预测蛋白质结构中的蛋白质结合位点

DOI:
10.1186/1471-2105-10-381
复制
发表时间:
2009-11-20
期刊:
影响因子:
3
通讯作者:
Wang X
Wang X
中科院分区:
生物学4区
文献类型:
--
作者:
Liu B;Wang X;Lin L;Tang B;Dong Q;Wang X

文献摘要

参考文献

被引文献

相似文献

预测两个相互作用的蛋白质之间的结合位点为蛋白质的功能提供了重要的线索。近年来对蛋白质结合位点预测的研究主要基于广为人知的机器学习技术,如人工神经网络、支持向量机、条件随机场等。然而,该方法的预测性能仍然很低,无法应用于实际。有必要探索新的算法、理论和特征来进一步提高性能。结果在本研究中,我们引入了一种新的机器学习模型隐马尔可夫支持向量机用于蛋白质结合位点预测。该模型将蛋白质结合位点预测作为基于最大边际准则的顺序标记任务。利用蛋白质序列和结构的共同特征,包括蛋白质序列轮廓和残基可达表面积,来训练隐马尔可夫支持向量机。在6个数据集上的测试表明,基于隐马尔可夫支持向量机的方法比人工神经网络、支持向量机和条件随机场等现有方法表现出更好的性能。该方法的运行时间比所比较的方法缩短了几个数量级。结论基于隐马尔可夫支持向量机的方法预测性能和计算效率的提高可归因于以下三个因素:首先,邻近残基标记之间的关系有助于蛋白质结合位点的预测。其次,核技巧在这一领域非常有利。第三,隐马尔可夫支持向量机的训练步骤复杂度与训练样本数量成线性关系。
BackgroundPredicting the binding sites between two interacting proteins provides important clues to the function of a protein. Recent research on protein binding site prediction has been mainly based on widely known machine learning techniques, such as artificial neural networks, support vector machines, conditional random field, etc. However, the prediction performance is still too low to be used in practice. It is necessary to explore new algorithms, theories and features to further improve the performance.ResultsIn this study, we introduce a novel machine learning model hidden Markov support vector machine for protein binding site prediction. The model treats the protein binding site prediction as a sequential labelling task based on the maximum margin criterion. Common features derived from protein sequences and structures, including protein sequence profile and residue accessible surface area, are used to train hidden Markov support vector machine. When tested on six data sets, the method based on hidden Markov support vector machine shows better performance than some state-of-the-art methods, including artificial neural networks, support vector machines and conditional random field. Furthermore, its running time is several orders of magnitude shorter than that of the compared methods.ConclusionThe improved prediction performance and computational efficiency of the method based on hidden Markov support vector machine can be attributed to the following three factors. Firstly, the relation between labels of neighbouring residues is useful for protein binding site prediction. Secondly, the kernel trick is very advantageous to this field. Thirdly, the complexity of the training step for hidden Markov support vector machine is linear with the number of training samples by using the cutting-plane algorithm.
DOI: 10.1016/s0097-8485(96)80004-0
发表时间: 1996-03-01
期刊: COMPUTERS & CHEMISTRY
影响因子: --
作者:
Gribskov, M;Robinson, NL
通讯作者: Robinson, NL
DOI: 10.1371/journal.pcbi.0020124
发表时间: 2006-09-29
影响因子: 4.3
作者:
Kim WK;Henschel A;Winter C;Schroeder M
通讯作者: Schroeder M
DOI: 10.1038/256705a0
发表时间: 1975-01-01
期刊: NATURE
影响因子: 64.8
作者:
CHOTHIA, C;JANIN, J
通讯作者: JANIN, J
DOI: 10.1093/protein/gzh020
发表时间: 2004-02-01
影响因子: 2.4
作者:
Koike, A;Takagi, T
通讯作者: Takagi, T
使用结构基序描述符进行基于序列的结合位点预测。
DOI: 10.1186/1471-2105-8-s4-s5
发表时间: 2007-05-22
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Henschel, Andreas;Winter, Christof;Kim, Wan Kyu;Schroeder, Michael
通讯作者: Schroeder, Michael