Prediction of protein solvent accessibility using fuzzy k-nearest neighbor method

Prediction of protein solvent accessibility using fuzzy k-nearest neighbor method
复制标题

DOI:
10.1093/bioinformatics/bti423
复制
发表时间:
2005-06-15
期刊:
影响因子:
5.8
通讯作者:
Lee, J
Lee, J
中科院分区:
生物学3区
文献类型:
--
作者:
Sim, J;Kim, SY;Lee, J

文献摘要

被引文献

相似文献

动机:氨基酸残基的溶剂可及性在三级结构预测中起着重要作用,特别是在一个查询蛋白与已知结构没有显著序列相似性的情况下。尽管近年来的研究有所改进,但溶剂可及性预测的准确性仍低于二级结构预测。k近邻法是一种简单但功能强大的分类算法,虽然它经常用于生物和医学数据的分类,但从未应用于溶剂可及性的预测。结果:将模糊k近邻法应用于溶剂可及性预测,以PSI-BLAST剖面作为特征向量,预测精度较高。通过序列聚类构建的ASTRAL SCOP参考数据集的留一交叉验证,我们的方法对3状态(埋藏/中间/暴露)预测(埋藏/中间阈值为9%,中间/暴露阈值为36%)的准确率达到64.1%,对2状态(埋藏/暴露)预测(埋藏/暴露阈值分别为0、5、16和25%)的准确率分别达到86.7、82.0、79.0和78.5%。在RS126数据集和229种蛋白质的基准数据集上,我们的方法也比其他方法略微提高了2-5%的准确性。
Motivation: The solvent accessibility of amino acid residues plays an important role in tertiary structure prediction, especially in the absence of significant sequence similarity of a query protein to those with known structures. The prediction of solvent accessibility is less accurate than secondary structure prediction in spite of improvements in recent researches. The k-nearest neighbor method, a simple but powerful classification algorithm, has never been applied to the prediction of solvent accessibility, although it has been used frequently for the classification of biological and medical data. Results: We applied the fuzzy k-nearest neighbor method to the solvent accessibility prediction, using PSI-BLAST profiles as feature vectors, and achieved high prediction accuracies. With leave-one-out cross-validation on the ASTRAL SCOP reference dataset constructed by sequence clustering, our method achieved 64.1 % accuracy for a 3-state (buried/intermediate/exposed) prediction (thresholds of 9% for buried/intermediate and 36% for intermediate/exposed) and 86.7, 82.0, 79.0 and 78.5% accuracies for 2-state (buried/exposed) predictions (thresholds of each 0, 5, 16 and 25% for buried/exposed), respectively. Our method also showed slightly better accuracies than other methods by about 2-5% on the RS126 dataset and a bench-marking dataset with 229 proteins.