SPINE-D: accurate prediction of short and long disordered regions by a single neural-network based method.

SPINE-D: accurate prediction of short and long disordered regions by a single neural-network based method.
复制标题

SPINE-D:通过基于单一神经网络的方法准确预测短和长无序区域

DOI:
10.1080/073911012010525022
复制
发表时间:
2012
影响因子:
4.4
通讯作者:
Zhou Y
Zhou Y
中科院分区:
生物学3区
文献类型:
--
作者:
Zhang T;Faraggi E;Xue B;Dunker AK;Uversky VN;Zhou Y

文献摘要

参考文献

被引文献

相似文献

蛋白质的短无序区和长无序区对不同的氨基酸残基具有不同的偏好。通常需要训练不同的方法来分别预测它们。在这项研究中,我们开发了一种称为SPINE-D的基于神经网络的技术,该技术首先进行三态预测(短无序区域和长无序区域中的有序残基和无序残基),然后将其还原为两态预测。SPINE-D在由Disprot注释的蛋白质和直接来自PDB的蛋白质的不同组合组成的各种集合上进行测试,所述PDB注释为在X射线确定的结构中通过缺失坐标的紊乱。虽然根据Disprot和X射线方法,无序注释是不同的,但SPINE-D的预测准确性和预测无序的能力相对独立于该方法是如何训练的,以及采用了什么类型的注释,但强烈依赖于测试集中短和长无序区域中有序和无序残基的相对群体的平衡。对于检测短无序区和长无序区中的残基具有大于85%的总体特异性,长无序区中的残基在具有56.5%有序残基的平衡测试数据集中以81%的灵敏度更容易预测,但在具有90%有序残基的测试数据集中更具挑战性(以65%的灵敏度)。与其他11种方法相比,SPINE-D产生了最高的曲线下面积(AUC),最高的基于残差的预测的马修斯相关系数,以及最低的均方误差在预测蛋白质的无序含量的329个蛋白质的独立测试集。特别地,SPINE-D在预测长无序区域中的无序残基方面与Meta预测因子相当,并且在短无序区域中具有上级优势。SPINE-D参加了CASP 9盲预测,根据官方排名,它是顶级服务器之一。此外,SPINE-D在几个案例研究中用于预测功能性分子识别基序。服务器和数据库可在http://sparks.informatics.iupui.edu/上获得。
Short and long disordered regions of proteins have different preference for different amino acid residues. Different methods often have to be trained to predict them separately. In this study, we developed a single neural-network-based technique called SPINE-D that makes a three-state prediction first (ordered residues and disordered residues in short and long disordered regions) and reduces it into a two-state prediction afterwards. SPINE-D was tested on various sets composed of different combinations of Disprot annotated proteins and proteins directly from the PDB annotated for disorder by missing coordinates in X-ray determined structures. While disorder annotations are different according to Disprot and X-ray approaches, SPINE-D's prediction accuracy and ability to predict disorder are relatively independent of how the method was trained and what type of annotation was employed but strongly depend on the balance in the relative populations of ordered and disordered residues in short and long disordered regions in the test set. With greater than 85% overall specificity for detecting residues in both short and long disordered regions, the residues in long disordered regions are easier to predict at 81% sensitivity in a balanced test dataset with 56.5% ordered residues but more challenging (at 65% sensitivity) in a test dataset with 90% ordered residues. Compared to eleven other methods, SPINE-D yields the highest area under the curve (AUC), the highest Mathews correlation coefficient for residue-based prediction, and the lowest mean square error in predicting disorder contents of proteins for an independent test set with 329 proteins. In particular, SPINE-D is comparable to a meta predictor in predicting disordered residues in long disordered regions and superior in short disordered regions. SPINE-D participated in CASP 9 blind prediction and is one of the top servers according to the official ranking. In addition, SPINE-D was examined for prediction of functional molecular recognition motifs in several case studies. The server and databases are available at http://sparks.informatics.iupui.edu/.
DOI: 10.1002/jcc.21968
发表时间: 2012-01-30
影响因子: 3
作者:
Faraggi, Eshel;Zhang, Tuo;Yang, Yuedong;Kurgan, Lukasz;Zhou, Yaoqi
通讯作者: Zhou, Yaoqi
DOI: 10.1016/j.sbi.2008.10.002
发表时间: 2008-12-01
影响因子: 6.8
作者:
Dunker, A. Keith;Silman, Israel;Sussman, Joel L.
通讯作者: Sussman, Joel L.
DOI: 10.1016/j.sbi.2008.12.004
发表时间: 2009-02
影响因子: 6.8
作者:
Eliezer, David
通讯作者: Eliezer, David
蛋白质障碍在多个敏感性和特异性水平上的预测。
DOI: 10.1186/1471-2164-9-s1-s9
发表时间: 2008
期刊: BMC GENOMICS
影响因子: 4.4
作者:
Hecker, Joshua;Yang, Jack Y.;Cheng, Jianlin
通讯作者: Cheng, Jianlin
DOI: 10.1093/bioinformatics/bti541
发表时间: 2005-08-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Dosztányi, Z;Csizmok, V;Simon, I
通讯作者: Simon, I