Sann: Solvent accessibility prediction of proteins by nearest neighbor method

Sann: Solvent accessibility prediction of proteins by nearest neighbor method
复制标题

DOI:
10.1002/prot.24074
复制
发表时间:
2012-07-01
影响因子:
2.9
通讯作者:
Lee, Jooyoung
Lee, Jooyoung
中科院分区:
生物学4区
文献类型:
--
作者:
Joo, Keehyoung;Lee, Sung Jong;Lee, Jooyoung

文献摘要

被引文献

相似文献

我们提出了一种预测蛋白质溶剂可及性的方法,该方法基于应用于序列图谱的最近邻方法。利用该方法可以得到连续的实值预报以及两状态和三状态的离散预报。该方法利用特征向量空间中距离度量的z-Score值来估计k近邻之间的相对贡献,以预测离散和连续的溶剂可及性。溶剂可及性数据库由5717个从双鱼剔除服务器中提取的蛋白质构建而成,截止点为25%的序列同源性。使用最优参数,预测精度(离散预测)为78.38%(阈值为25%的二态预测)、65.1%(阈值为9%和36%的三态预测)和0.676的皮尔逊相关系数(预测与真实RSA连续预测之间的相关系数)。预测精度为80.89%(阈值为25%的二态预测)、67.58%(三态预测)、皮尔逊相关系数0.727(连续预测),平均绝对误差为0.148。我们还研究了数据库大小增加对预测精度的影响,其中随着数据库大小的增加,预测精度会进一步提高。Sann网络服务器可在Proteins 2012上获得;(C)2012 Wiley期刊,Inc.
We present a method to predict the solvent accessibility of proteins which is based on a nearest neighbor method applied to the sequence profiles. Using the method, continuous real-value prediction as well as two-state and three-state discrete predictions can be obtained. The method utilizes the z-score value of the distance measure in the feature vector space to estimate the relative contribution among the k-nearest neighbors for prediction of the discrete and continuous solvent accessibility. The Solvent accessibility database is constructed from 5717 proteins extracted from PISCES culling server with the cutoff of 25% sequence identities. Using optimal parameters, the prediction accuracies (for discrete predictions) of 78.38% (two-state prediction with the threshold of 25%), 65.1% (three-state prediction with the thresholds of 9 and 36%), and the Pearson correlation coefficient (between the predicted and true RSA's for continuous prediction) of 0.676 are achieved An independent benchmark test was performed with the CASP8 targets where we find that the proposed method outperforms existing methods. The prediction accuracies are 80.89% (for two state prediction with the threshold of 25%), 67.58% (three-state prediction), and the Pearson correlation coefficient of 0.727 (for continuous prediction) with mean absolute error of 0.148. We have also investigated the effect of increasing database sizes on the prediction accuracy, where additional improvement in the accuracy is observed as the database size increases. The SANN web server is available at .Proteins 2012; (c) 2012 Wiley Periodicals, Inc.