Reaching alignment-profile-based accuracy in predicting protein secondary and tertiary structural properties without alignment.

Reaching alignment-profile-based accuracy in predicting protein secondary and tertiary structural properties without alignment.
复制标题

DOI:
10.1038/s41598-022-11684-w
复制
发表时间:
2022-05-09
期刊:
影响因子:
4.6
通讯作者:
--
中科院分区:
综合性期刊3区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

蛋白质语言模型已经作为多序列比对的替代方案出现,用于丰富序列信息和改善下游预测任务,如生物物理、结构和功能特性。在这里,我们展示了一种名为SPOT-1D-LM的方法,该方法将传统的单热编码与来自两个不同语言模型(ProtTrans和ESM-1b)的嵌入相结合,在预测蛋白质1D二级和三级结构属性(包括所有六个测试集(TEST2018、TEST2020、Neff1-2020、CASP12-FM、CASP13-FM和CASP14-FM)的主干扭转角、溶剂可及性和联系电话)方面,与基于单序列的技术相比,精度有了飞跃。更重要的是,对于那些具有同源序列的蛋白质,它具有与基于图谱的方法相当的性能。例如,SPOT-1D-LM对TEST2018和TEST2020蛋白质的三态二级结构(SS3)预测的准确率分别为86.7%和79.8%,而基于单序列的方法SPOT-1D-Single和基于Profile的方法SPOT-1D分别为74.3%和73.4%,86.2%和80.5%。对于没有同源序列(Neff1-2020)的蛋白质,SPOT-1D-LM的SS3为80.41%,分别比SPOT-1D-Single和Spot-1D高3.8%和8.3%。由于其快速的性能,Spot-1D-LM有望用于全基因组分析。此外,在没有序列比对的情况下,对二级和三级结构属性(如主干角度和溶剂可及性)的高精度预测表明,可以在没有同源序列的情况下对蛋白质结构进行高精度预测,这是后AlphaFold2时代的剩余障碍。
Protein language models have emerged as an alternative to multiple sequence alignment for enriching sequence information and improving downstream prediction tasks such as biophysical, structural, and functional properties. Here we show that a method called SPOT-1D-LM combines traditional one-hot encoding with the embeddings from two different language models (ProtTrans and ESM-1b) for the input and yields a leap in accuracy over single-sequence-based techniques in predicting protein 1D secondary and tertiary structural properties, including backbone torsion angles, solvent accessibility and contact numbers for all six test sets (TEST2018, TEST2020, Neff1-2020, CASP12-FM, CASP13-FM and CASP14-FM). More significantly, it has a performance comparable to profile-based methods for those proteins with homologous sequences. For example, the accuracy for three-state secondary structure (SS3) prediction for TEST2018 and TEST2020 proteins are 86.7% and 79.8% by SPOT-1D-LM, compared to 74.3% and 73.4% by the single-sequence-based method SPOT-1D-Single and 86.2% and 80.5% by the profile-based method SPOT-1D, respectively. For proteins without homologous sequences (Neff1-2020) SS3 is 80.41% by SPOT-1D-LM which is 3.8% and 8.3% higher than SPOT-1D-Single and SPOT-1D, respectively. SPOT-1D-LM is expected to be useful for genome-wide analysis given its fast performance. Moreover, high-accuracy prediction of both secondary and tertiary structural properties such as backbone angles and solvent accessibility without sequence alignment suggests that highly accurate prediction of protein structures may be made without homologous sequences, the remaining obstacle in the post AlphaFold2 era.
DOI: 10.1038/s41586-021-03819-2
发表时间: 2021-08
期刊: Nature
影响因子: 64.8
作者:
Jumper J;Evans R;Pritzel A;Green T;Figurnov M;Ronneberger O;Tunyasuvunakool K;Bates R;Žídek A;Potapenko A;Bridgland A;Meyer C;Kohl SAA;Ballard AJ;Cowie A;Romera-Paredes B;Nikolov S;Jain R;Adler J;Back T;Petersen S;Reiman D;Clancy E;Zielinski M;Steinegger M;Pacholska M;Berghammer T;Bodenstein S;Silver D;Vinyals O;Senior AW;Kavukcuoglu K;Kohli P;Hassabis D
通讯作者: Hassabis D
DOI: 10.1016/j.gpb.2019.01.004
发表时间: 2019-12-01
影响因子: 9.5
作者:
Hanson, Jack;Paliwal, Kuldip K.;Zhou, Yaoqi
通讯作者: Zhou, Yaoqi
DOI: 10.1093/bioinformatics/btv665
发表时间: 2016-03-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Heffernan, Rhys;Dehzangi, Abdollah;Yang, Yuedong
通讯作者: Yang, Yuedong
DOI: 10.1002/prot.25674
发表时间: 2019-06-01
影响因子: 2.9
作者:
Klausen, Michael Schantz;Jespersen, Martin Closter;Marcatili, Paolo
通讯作者: Marcatili, Paolo
DOI: 10.1002/0471250953.bi0301s42
发表时间: 2013-06
影响因子: --
作者:
Pearson, William R
通讯作者: Pearson, William R