Predicting secondary structures, contact numbers, and residue-wise contact orders of native protein structures from amino acid sequences using critical random networks.

Predicting secondary structures, contact numbers, and residue-wise contact orders of native protein structures from amino acid sequences using critical random networks.
复制标题

DOI:
10.2142/biophysics.1.67
复制
发表时间:
2005
期刊:
Biophysics (Nagoya-shi, Japan)
影响因子:
--
通讯作者:
Nishikawa K
Nishikawa K
中科院分区:
其他
文献类型:
--
作者:
Kinjo AR;Nishikawa K

文献摘要

被引文献

相似文献

预测蛋白质的一维结构,如二级结构和接触数,对于预测三维结构是有用的,对于理解序列-结构关系也很重要。在这里,我们提出了一种新的机器学习方法,临界随机网络(CRN),用于预测一维结构,并将其应用于二级结构(SS)、接触数(CN)和残基接触顺序(RWCO)的预测。该方法对SS的Q3正确率为77.8%,对CN和RWCO的相关系数分别为0.726和0.601。SS预报的精度与其他先进方法相当,CN预报的精度比以前的方法有了显著的提高。我们给出了基于临界随机网络的预测方案的详细公式,并考察了预测精度的上下文相关性。为了研究非线性和多体效应,我们将基于CRN的方法与基于位置特定评分矩阵的纯线性方法进行了比较。虽然不优于基于CRNS的方法,但线性方法获得的惊人的高精度突显了从氨基酸序列中提取高阶结构特征的难度,这些结构特征超出了特定位置评分矩阵所提供的信息。
Predictions of one-dimensional protein structures such as secondary structures and contact numbers are useful for predicting three-dimensional structure and important for understanding the sequence-structure relationship. Here we present a new machine-learning method, critical random networks (CRNs), for predicting one-dimensional structures, and apply it, with position-specific scoring matrices, to the prediction of secondary structures (SS), contact numbers (CN), and residue-wise contact orders (RWCO). The present method achieves, on average, Q3 accuracy of 77.8% for SS, and correlation coefficients of 0.726 and 0.601 for CN and RWCO, respectively. The accuracy of the SS prediction is comparable to that obtained with other state-of-the-art methods, and accuracy of the CN prediction is a significant improvement over that with previous methods. We give a detailed formulation of the critical random networks-based prediction scheme, and examine the context-dependence of prediction accuracies. In order to study the nonlinear and multi-body effects, we compare the CRNs-based method with a purely linear method based on position-specific scoring matrices. Although not superior to the CRNs-based method, the surprisingly good accuracy achieved by the linear method highlights the difficulty in extracting structural features of higher order from an amino acid sequence beyond the information provided by the position-specific scoring matrices.