NetSurfP-2.0: Improved prediction of protein structural features by integrated deep learning

NetSurfP-2.0: Improved prediction of protein structural features by integrated deep learning
复制标题

DOI:
10.1002/prot.25674
复制
发表时间:
2019-06-01
影响因子:
2.9
通讯作者:
Marcatili, Paolo
Marcatili, Paolo
中科院分区:
生物学4区
文献类型:
--
作者:
Klausen, Michael Schantz;Jespersen, Martin Closter;Marcatili, Paolo

文献摘要

被引文献

相似文献

从一级序列预测蛋白质局部结构特征的能力对于在缺乏实验结构信息的情况下揭示其功能至关重要。两个主要因素影响潜在预测工具的效用:它们的准确性必须能够提取感兴趣蛋白质的可靠结构信息,并且它们的运行时间必须低,以跟上以不断增加的速度生成的测序数据。在这里,我们提出了NetSurfP-2.0,一种新的工具,可以预测最重要的局部结构特征,具有前所未有的准确性和运行时间。NetSurfP-2.0是基于序列的,使用由在已解决的蛋白质结构上训练的卷积和长短期记忆神经网络组成的架构。使用一个单一的集成模型,NetSurfP-2.0预测溶剂的可及性,二级结构,结构紊乱,和骨干二面角的输入序列的每个残基。我们在几个独立的测试数据集上评估了NetSurfP-2.0的准确性,发现它能够始终如一地为其每个输出特征产生最先进的预测。我们观察到的预测和实验数据之间的相关性为80%的溶剂可及性,和85%的精度二级结构3级预测。除了提高准确性外,还优化了处理时间,允许在不到2小时内预测1000多个蛋白质,并在不到1天内完成蛋白质组。
The ability to predict local structural features of a protein from the primary sequence is of paramount importance for unraveling its function in absence of experimental structural information. Two main factors affect the utility of potential prediction tools: their accuracy must enable extraction of reliable structural information on the proteins of interest, and their runtime must be low to keep pace with sequencing data being generated at a constantly increasing speed. Here, we present NetSurfP-2.0, a novel tool that can predict the most important local structural features with unprecedented accuracy and runtime. NetSurfP-2.0 is sequence-based and uses an architecture composed of convolutional and long short-term memory neural networks trained on solved protein structures. Using a single integrated model, NetSurfP-2.0 predicts solvent accessibility, secondary structure, structural disorder, and backbone dihedral angles for each residue of the input sequences. We assessed the accuracy of NetSurfP-2.0 on several independent test datasets and found it to consistently produce state-of-the-art predictions for each of its output features. We observe a correlation of 80% between predictions and experimental data for solvent accessibility, and a precision of 85% on secondary structure 3-class predictions. In addition to improved accuracy, the processing time has been optimized to allow predicting more than 1000 proteins in less than 2 hours, and complete proteomes in less than 1 day.