Prediction of Carbohydrate-Binding Proteins from Sequences Using Support Vector Machines

Prediction of Carbohydrate-Binding Proteins from Sequences Using Support Vector Machines
复制标题

DOI:
10.1155/2010/289301
复制
发表时间:
2010-09
影响因子:
--
通讯作者:
Seizi Someya;Masanori Kakuta;Mizuki Morita;K. Sumikoshi;Cao Wei;Zhenyi Ge;Osamu Hirose;Shugo Nakamura;T. Terada;K. Shimizu
Seizi Someya;Masanori Kakuta;Mizuki Morita;K. Sumikoshi;Cao Wei;Zhenyi Ge;Osamu Hirose;Shugo Nakamura;T. Terada;K. Shimizu
中科院分区:
--
文献类型:
--
作者:
Seizi Someya;Masanori Kakuta;Mizuki Morita;K. Sumikoshi;Cao Wei;Zhenyi Ge;Osamu Hirose;Shugo Nakamura;T. Terada;K. Shimizu

文献摘要

相似文献

碳水化合物结合蛋白是可以与糖链相互作用但不能修饰糖链的蛋白质。它们参与许多生理功能,我们已经开发出一种根据它们的氨基酸序列来预测它们的方法。我们的方法是基于支持向量机的。我们首先澄清了碳水化合物结合蛋白的定义,然后构建了用于训练支持向量机的正数据集和负数据集。通过对这些数据集进行留一检验,我们的方法提供了接收器工作特征(ROC)曲线下0.92的面积。我们还研究了两种氨基酸分组方法,它们能够有效地学习序列模式,并评估了这些方法的性能。当我们将我们的方法与基于同源性的预测方法相结合应用于带注释的人类基因组数据库H-invDB时,我们发现预测的真阳性率得到了提高。
Carbohydrate-binding proteins are proteins that can interact with sugar chains but do not modify them. They are involved in many physiological functions, and we have developed a method for predicting them from their amino acid sequences. Our method is based on support vector machines (SVMs). We first clarified the definition of carbohydrate-binding proteins and then constructed positive and negative datasets with which the SVMs were trained. By applying the leave-one-out test to these datasets, our method delivered 0.92 of the area under the receiver operating characteristic (ROC) curve. We also examined two amino acid grouping methods that enable effective learning of sequence patterns and evaluated the performance of these methods. When we applied our method in combination with the homology-based prediction method to the annotated human genome database, H-invDB, we found that the true positive rate of prediction was improved.