Predicting the errors of predicted local backbone angles and non-local solvent-accessibilities of proteins by deep neural networks

Predicting the errors of predicted local backbone angles and non-local solvent-accessibilities of proteins by deep neural networks
复制标题

DOI:
10.1093/bioinformatics/btw549
复制
发表时间:
2016-12-15
期刊:
影响因子:
5.8
通讯作者:
Zhou, Yaoqi
Zhou, Yaoqi
中科院分区:
生物学3区
文献类型:
--
作者:
Gao, Jianzhao;Yang, Yuedong;Zhou, Yaoqi

文献摘要

被引文献

相似文献

动机:蛋白质的主链结构和溶剂可及表面积受益于连续的实值预测,因为它消除了定义不同二级结构和溶剂可及状态之间边界的任意性。然而,缺乏预测值的置信度限制了它们的应用。在这里,我们研究了是否可以通过采用深度神经网络对预测的主干扭转角、基于钙原子的角度和扭转角、溶剂可及性、接触次数和半球暴露量的绝对误差进行合理的预测。结果:我们发现基于角度的误差可以最准确地预测,预测误差和实际误差之间的 Spearman 相关系数 (SPC) 约为 0.6。其次是溶剂可及性(SPC 类似于 0.5)。基于接触的结构特性的误差是最难预测的(SPC 在 0.2 到 0.3 之间)。我们表明,预测误差是比基于二级结构和氨基酸残基类型的平均误差明显更好的误差指标。我们进一步证明了预测误差在模型质量评估中的有用性。这些误差或置信度指标预计可用于蛋白质结构的预测、评估和细化。可用性和实施​​:该方法可作为 SPIDER2 包的一部分在 http://sparks-lab.org 上获取。联系方式:yuedong.yang@griffith.edu.au 或 yaoqi.zhou@griffith.edu.au 补充信息:补充数据可在 生物信息学在线。
Motivation: Backbone structures and solvent accessible surface area of proteins are benefited from continuous real value prediction because it removes the arbitrariness of defining boundary between different secondary-structure and solvent-accessibility states. However, lacking the confidence score for predicted values has limited their applications. Here we investigated whether or not we can make a reasonable prediction of absolute errors for predicted backbone torsion angles, Ca-atom-based angles and torsion angles, solvent accessibility, contact numbers and half-sphere exposures by employing deep neural networks.Results: We found that angle-based errors can be predicted most accurately with Spearman correlation coefficient (SPC) between predicted and actual errors at about 0.6. This is followed by solvent accessibility (SPC similar to 0.5). The errors on contact-based structural properties are most difficult to predict (SPC between 0.2 and 0.3). We showed that predicted errors are significantly better error indicators than the average errors based on secondary-structure and amino-acid residue types. We further demonstrated the usefulness of predicted errors in model quality assessment. These error or confidence indictors are expected to be useful for prediction, assessment, and refinement of protein structures.Availability and Implementation: The method is available at http://sparks-lab.org as a part of SPIDER2 package.Contact: yuedong.yang@griffith.edu.au or yaoqi.zhou@griffith.edu.auSupplementary information: Supplementary data are available at Bioinformatics online.