Development of a sugar-binding residue prediction system from protein sequences using support vector machine

Development of a sugar-binding residue prediction system from protein sequences using support vector machine
复制标题

DOI:
10.1016/j.compbiolchem.2016.10.009
复制
发表时间:
2017-02-01
影响因子:
3.1
通讯作者:
Shimizu, Kentaro
Shimizu, Kentaro
中科院分区:
生物学3区
文献类型:
--
作者:
Banno, Masaki;Komiyama, Yusuke;Shimizu, Kentaro

文献摘要

被引文献

相似文献

已经提出了几种使用机器学习算法进行蛋白质-糖结合位点预测的方法。然而,它们不能有效地学习由蛋白质和糖之间的各种相互作用引起的结合位点残基的各种性质。在这项研究中,我们将糖分为酸性和非酸性糖,并表明它们的结合位点具有不同的氨基酸出现频率。通过使用这一结果,我们开发了糖结合残基预测专用于两类糖:酸性糖结合预测和非酸性糖结合预测。我们还开发了一个组合预测器,它结合了两个预测器的结果。我们表明,当已知糖是酸性糖时,酸性糖结合预测器将获得最佳性能,并表明当已知糖是非酸性糖或未知是这两种糖中的任何一种时,组合预测器将获得最佳性能。我们的方法仅使用氨基酸序列进行预测。采用支持向量机作为机器学习算法,采用位置特定迭代基本局部比对搜索工具创建的位置特定评分矩阵作为特征向量。我们使用五重交叉验证评估了预测因子的性能。我们已经推出了我们的系统,作为GitHub存储库(https://doi.org/10.5281izenodo.61513)上的开源免费工具。(C)2016年6月,作者。爱思唯尔有限公司出版
Several methods have been proposed for protein-sugar binding site prediction using machine learning algorithms. However, they are not effective to learn various properties of binding site residues caused by various interactions between proteins and sugars. In this study, we classified sugars into acidic and nonacidic sugars and showed that their binding sites have different amino acid occurrence frequencies. By using this result, we developed sugar-binding residue predictors dedicated to the two classes of sugars: an acid sugar binding predictor and a nonacidic sugar binding predictor. We also developed a combination predictor which combines the results of the two predictors. We showed that when a sugar is known to be an acidic sugar, the acidic sugar binding predictor achieves the best performance, and showed that when a sugar is known to be a nonacidic sugar or is not known to be either of the two classes, the combination predictor achieves the best performance. Our method uses only amino acid sequences for prediction. Support vector machine was used as a machine learning algorithm and the position-specific scoring matrix created by the position-specific iterative basic local alignment search tool was used as the feature vector. We evaluated the performance of the predictors using five-fold cross-validation. We have launched our system, as an open source freeware tool on the GitHub repository (https://doi.org/10.5281izenodo.61513). (C) 2016 The Authors. Published by Elsevier Ltd.