Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function.

Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function.
复制标题

DOI:
10.1093/bioinformatics/btad208
复制
发表时间:
2023-06-30
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

无数的基因组和转录组测序项目已经产生了数百万个蛋白质序列。然而,通过实验确定蛋白质的功能仍然是一个耗时、低通量和昂贵的过程,导致蛋白质序列-功能差距很大。因此,开发准确预测蛋白质功能的计算方法来填补这一空白是很重要的。尽管已经开发了许多利用蛋白质序列作为输入来预测功能的方法,但由于直到最近还缺乏针对大多数蛋白质的准确的蛋白质结构,因此利用蛋白质结构来预测蛋白质功能的方法要少得多。我们开发了TransFun-a方法,使用基于变压器的蛋白质语言模型和3D等变图神经网络来从蛋白质序列和结构中提取信息来预测蛋白质功能。该算法通过迁移学习利用预先训练好的蛋白质语言模型(ESM)从蛋白质序列中提取特征嵌入,并通过等变图神经网络将其与AlphaFold2预测的蛋白质结构相结合。以CAFA3测试数据集和新的测试数据集为基准,TransFun的性能优于几种最先进的方法,表明语言模型和3D等变图神经网络是利用蛋白质序列和结构来提高蛋白质功能预测的有效方法。结合TransFun预测和基于序列相似性的预测可以进一步提高预测精度。TransFun的源代码可以在https://github.com/jianlin-cheng/TransFun.上找到
Millions of protein sequences have been generated by numerous genome and transcriptome sequencing projects. However, experimentally determining the function of the proteins is still a time consuming, low-throughput, and expensive process, leading to a large protein sequence-function gap. Therefore, it is important to develop computational methods to accurately predict protein function to fill the gap. Even though many methods have been developed to use protein sequences as input to predict function, much fewer methods leverage protein structures in protein function prediction because there was lack of accurate protein structures for most proteins until recently. We developed TransFun—a method using a transformer-based protein language model and 3D-equivariant graph neural networks to distill information from both protein sequences and structures to predict protein function. It extracts feature embeddings from protein sequences using a pre-trained protein language model (ESM) via transfer learning and combines them with structures of proteins predicted by AlphaFold2 through equivariant graph neural networks. Benchmarked on the CAFA3 test dataset and a new test dataset, TransFun outperforms several state-of-the-art methods, indicating that the language model and 3D-equivariant graph neural networks are effective methods to leverage protein sequences and structures to improve protein function prediction. Combining TransFun predictions and sequence similarity-based predictions can further increase prediction accuracy. The source code of TransFun is available at https://github.com/jianlin-cheng/TransFun.
DOI: 10.1038/s41586-021-03819-2
发表时间: 2021-08
期刊: Nature
影响因子: 64.8
作者:
Jumper J;Evans R;Pritzel A;Green T;Figurnov M;Ronneberger O;Tunyasuvunakool K;Bates R;Žídek A;Potapenko A;Bridgland A;Meyer C;Kohl SAA;Ballard AJ;Cowie A;Romera-Paredes B;Nikolov S;Jain R;Adler J;Back T;Petersen S;Reiman D;Clancy E;Zielinski M;Steinegger M;Pacholska M;Berghammer T;Bodenstein S;Silver D;Vinyals O;Senior AW;Kavukcuoglu K;Kohli P;Hassabis D
通讯作者: Hassabis D
DOI: 10.1093/bioinformatics/btt228
发表时间: 2013-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Clark WT;Radivojac P
通讯作者: Radivojac P
DOI: 10.1093/bib/bbab502
发表时间: 2022-01-17
影响因子: 9.5
作者:
Lai, Boqiao;Xu, Jinbo
通讯作者: Xu, Jinbo
DOI: 10.1109/tpami.2021.3095381
发表时间: 2022-10-01
影响因子: 23.6
作者:
Elnaggar, Ahmed;Heinzinger, Michael;Rost, Burkhard
通讯作者: Rost, Burkhard
DOI: 10.1093/nar/gki414
发表时间: 2005-07-01
影响因子: 14.9
作者:
Laskowski RA;Watson JD;Thornton JM
通讯作者: Thornton JM