SETH predicts nuances of residue disorder from protein embeddings.

SETH predicts nuances of residue disorder from protein embeddings.
复制标题

塞思预测蛋白质嵌入的残基障碍的细微差别。

DOI:
10.3389/fbinf.2022.1019597
复制
发表时间:
2022
期刊:
FRONTIERS IN BIOINFORMATICS
影响因子:
--
通讯作者:
Rost, Burkhard
Rost, Burkhard
中科院分区:
其他
文献类型:
--
作者:
Ilzhofer, Dagmar;Heinzinger, Michael;Rost, Burkhard

文献摘要

参考文献

被引文献

相似文献

自从UniProt的AlphaFold2结果发布以来,数百万蛋白质三维结构的预测只需点击几下即可。然而,许多蛋白质具有所谓的内在无序区(IDR),这些区域在孤立的情况下并不采用独特的结构。这些IDR与几种疾病有关,包括阿尔茨海默病。我们发现,最近的三个障碍措施AlphaFold2预测(pLDDT,“实验解决”的预测和“相对溶剂可及性”)在一定程度上与IDR相关。然而,专家方法通过将复杂的机器学习模型与专家制作的输入特征和来自多序列比对(MSA)的进化信息相结合,可以更可靠地预测IDR。MSA并不总是可用的,特别是对于IDR,并且生成时计算成本很高,限制了相关工具的可伸缩性。在这里,我们提出了一种新的方法SETH,该方法从蛋白质语言模型ProtT5生成的嵌入中预测残基紊乱,该模型明确地只使用单个序列作为输入。因此,我们的方法依赖于相对较浅的卷积神经网络,性能优于更复杂的解决方案,同时速度更快,允许在消费级PC上使用NVIDIA GeForce RTX 3060在大约1小时内创建人类蛋白质组预测。经过连续混乱量表(CheZOD评分)的训练,我们的方法捕捉到了混乱中的微妙变化,从而提供了超出大多数方法二元分类的重要信息。高性能与速度相结合表明,SETH对整个蛋白质组的细微差别的无序预测捕捉了生物体进化的各个方面。此外,SETH还可以用于过滤出可能具有低质量AlphaFold2 3D结构的区域或蛋白质,以优先运行大型数据集的计算密集型预测。SETH可在https://github.com/Rostlab/SETH上免费公开获取。
Predictions for millions of protein three-dimensional structures are only a few clicks away since the release of AlphaFold2 results for UniProt. However, many proteins have so-called intrinsically disordered regions (IDRs) that do not adopt unique structures in isolation. These IDRs are associated with several diseases, including Alzheimer’s Disease. We showed that three recent disorder measures of AlphaFold2 predictions (pLDDT, “experimentally resolved” prediction and “relative solvent accessibility”) correlated to some extent with IDRs. However, expert methods predict IDRs more reliably by combining complex machine learning models with expert-crafted input features and evolutionary information from multiple sequence alignments (MSAs). MSAs are not always available, especially for IDRs, and are computationally expensive to generate, limiting the scalability of the associated tools. Here, we present the novel method SETH that predicts residue disorder from embeddings generated by the protein Language Model ProtT5, which explicitly only uses single sequences as input. Thereby, our method, relying on a relatively shallow convolutional neural network, outperformed much more complex solutions while being much faster, allowing to create predictions for the human proteome in about 1 hour on a consumer-grade PC with one NVIDIA GeForce RTX 3060. Trained on a continuous disorder scale (CheZOD scores), our method captured subtle variations in disorder, thereby providing important information beyond the binary classification of most methods. High performance paired with speed revealed that SETH’s nuanced disorder predictions for entire proteomes capture aspects of the evolution of organisms. Additionally, SETH could also be used to filter out regions or proteins with probable low-quality AlphaFold2 3D structures to prioritize running the compute-intensive predictions for large data sets. SETH is freely publicly available at: https://github.com/Rostlab/SETH.
DOI: 10.1007/s00239-021-10022-4
发表时间: 2021-10
影响因子: 3.9
作者:
Marot-Lassauzaie V;Goldberg T;Armenteros JJA;Nielsen H;Rost B
通讯作者: Rost B