Alignment-free estimation of sequence conservation for identifying functional sites using protein sequence embeddings.

Alignment-free estimation of sequence conservation for identifying functional sites using protein sequence embeddings.
复制标题

DOI:
10.1093/bib/bbac599
复制
发表时间:
2023-01-19
影响因子:
9.5
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

蛋白质语言建模是生物信息学中一种新兴的深度学习方法,在结构预测、蛋白质设计等领域有着广泛的应用。然而,功能位点预测在估计序列保守性方面的应用还没有被系统地探索。在这里,我们提出了一种使用从蛋白质语言模型生成的序列嵌入来估计序列保守性的无比对估计方法。跨公开可用的蛋白质语言模型的综合基准表明,ESM2模型为保守估计提供了最佳的性能与计算成本比。将我们的方法应用于全长蛋白质序列,我们证明了基于嵌入的方法对保守元件的顺序不敏感--可以在一次运行中计算多结构域蛋白质的保守分数,而不需要分离单独的结构域。我们的方法还可以在快速进化的序列区域(如结构域插入)中识别保守的功能位点,这一点我们通过在蛋白激酶的可变插入片段中识别保守的磷酸化基序来证明。总体而言,基于嵌入的保守性分析是一种广泛适用的方法,可以识别任何全长蛋白质序列中的潜在功能位点,并以不比对的方式估计保守性。要在您感兴趣的蛋白质序列上运行此程序,请尝试我们的脚本,网址为https://github.com/esbgkannan/kibby.
Protein language modeling is a fast-emerging deep learning method in bioinformatics with diverse applications such as structure prediction and protein design. However, application toward estimating sequence conservation for functional site prediction has not been systematically explored. Here, we present a method for the alignment-free estimation of sequence conservation using sequence embeddings generated from protein language models. Comprehensive benchmarks across publicly available protein language models reveal that ESM2 models provide the best performance to computational cost ratio for conservation estimation. Applying our method to full-length protein sequences, we demonstrate that embedding-based methods are not sensitive to the order of conserved elements—conservation scores can be calculated for multidomain proteins in a single run, without the need to separate individual domains. Our method can also identify conserved functional sites within fast-evolving sequence regions (such as domain inserts), which we demonstrate through the identification of conserved phosphorylation motifs in variable insert segments in protein kinases. Overall, embedding-based conservation analysis is a broadly applicable method for identifying potential functional sites in any full-length protein sequence and estimating conservation in an alignment-free manner. To run this on your protein sequence of interest, try our scripts at https://github.com/esbgkannan/kibby.
DOI: 10.1093/nar/gkaa913
发表时间: 2021-01-08
影响因子: 14.9
作者:
Mistry J;Chuguransky S;Williams L;Qureshi M;Salazar GA;Sonnhammer ELL;Tosatto SCE;Paladin L;Raj S;Richardson LJ;Finn RD;Bateman A
通讯作者: Bateman A
DOI: 10.1073/pnas.89.22.10915
发表时间: 1992-11-15
影响因子: 11.1
作者:
HENIKOFF, S;HENIKOFF, JG
通讯作者: HENIKOFF, JG
DOI: 10.1007/s00439-021-02411-y
发表时间: 2022-10
期刊: Human genetics
影响因子: 5.3
作者:
Marquet C;Heinzinger M;Olenyi T;Dallago C;Erckert K;Bernhofer M;Nechaev D;Rost B
通讯作者: Rost B
DOI: 10.1038/s41586-021-03819-2
发表时间: 2021-08
期刊: Nature
影响因子: 64.8
作者:
Jumper J;Evans R;Pritzel A;Green T;Figurnov M;Ronneberger O;Tunyasuvunakool K;Bates R;Žídek A;Potapenko A;Bridgland A;Meyer C;Kohl SAA;Ballard AJ;Cowie A;Romera-Paredes B;Nikolov S;Jain R;Adler J;Back T;Petersen S;Reiman D;Clancy E;Zielinski M;Steinegger M;Pacholska M;Berghammer T;Bodenstein S;Silver D;Vinyals O;Senior AW;Kavukcuoglu K;Kohli P;Hassabis D
通讯作者: Hassabis D
DOI: 10.1074/jbc.275.21.16219
发表时间: 2000-05-26
影响因子: 4.8
作者:
Kovalenko, M;Denner, K;Östman, A
通讯作者: Östman, A