Large language models generate functional protein sequences across diverse families.

Large language models generate functional protein sequences across diverse families.
复制标题

DOI:
10.1038/s41587-022-01618-2
复制
发表时间:
2023-08
影响因子:
46.9
通讯作者:
Naik, Nikhil
Naik, Nikhil
中科院分区:
工程技术1区
文献类型:
--
作者:
Madani, Ali;Ben Krause, Ben;Greene, Eric R.;Subramanian, Subu;Mohr, Benjamin P.;Holton, James M.;Olmos, Jose Luis;Xiong, Caiming;Sun, Zachary Z. Z.;Socher, Richard;Fraser, James S.;Naik, Nikhil

文献摘要

参考文献

被引文献

相似文献

利用人工智能生成人工蛋白质序列可以为生物医学和环境挑战提供突破性的解决方案。将氨基酸序列视为一种语言,我们证明了基于深度学习的语言模型可以在大型蛋白质家族中生成功能性人工蛋白质序列,类似于在不同主题上生成语法和语义正确的自然语言句子。我们的蛋白质语言模型通过简单的学习来预测来自数千个蛋白质家族的超过2.8亿个蛋白质序列的下一个氨基酸,而无需明确的生物物理建模。我们实验评估模型生成的人工蛋白微调到五种不同的抗菌溶菌酶家族。人工蛋白的活性和催化效率与典型的天然溶菌酶(包括蛋清溶菌酶)相似,但与任何已知的自然进化蛋白的同源性低至31.4%。酶活性人工蛋白的x射线晶体结构概括了天然蛋白中活性位点残基的保守折叠和定位。我们通过准确预测人工choris酸突变酶和苹果酸脱氢酶蛋白的功能,展示了我们的语言模型适应不同蛋白质家族的能力。这些结果表明,神经语言模型成功地完成了跨蛋白质家族的人工蛋白质生成,可能被证明是一种捷径进化的工具。
Generating artificial protein sequences using artificial intelligence could enable breakthrough solutions for biomedical and environmental challenges. Viewing amino acid sequences as a language, we demonstrate that a deep learning-based language model can generate functional artificial protein sequences across large protein families, akin to generating grammatically and semantically correct natural language sentences on diverse topics. Our protein language model is trained by simply learning to predict the next amino acid for over 280 million protein sequences from thousands of protein families, without explicit biophysical modeling. We experimentally evaluate model-generated artificial proteins fine-tuned to five distinct antibacterial lysozyme families. Artificial proteins show similar activities and catalytic efficiencies as representative natural lysozymes, including hen egg white lysozyme, while maintaining activity with as low as 31.4% identity to any known naturally-evolved protein. The X-ray crystal structure of an enzymatically active artificial protein recapitulates the conserved fold and positioning of active site residues found in natural proteins. We show our language model’s ability to be adapted to different protein families by accurately predicting the functionality of artificial chorismate mutase and malate dehydrogenase proteins. These results indicate that neural language models successfully perform artificial protein generation across protein families and may prove to be a tool to shortcut evolution.
DOI: 10.1093/nar/gkr1178
发表时间: 2012-01
影响因子: 14.9
作者:
Federhen S
通讯作者: Federhen S
DOI: 10.1109/tpami.2021.3095381
发表时间: 2022-10-01
影响因子: 23.6
作者:
Elnaggar, Ahmed;Heinzinger, Michael;Rost, Burkhard
通讯作者: Rost, Burkhard
DOI: 10.1093/nar/gkl929
发表时间: 2007-01-01
影响因子: 14.9
作者:
Bairoch, Amos;Bougueleret, Lydie;Zhang, Jian
通讯作者: Zhang, Jian
DOI: 10.1002/prot.22934
发表时间: 2011-04-01
影响因子: 2.9
作者:
Balakrishnan, Sivaraman;Kamisetty, Hetunandan;Langmead, Christopher James
通讯作者: Langmead, Christopher James
用于协同进化序列分析的Evcouplings Python框架。
DOI: 10.1093/bioinformatics/bty862
发表时间: 2019-05-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Hopf TA;Green AG;Schubert B;Mersmann S;Schärfe CPI;Ingraham JB;Toth-Petroczy A;Brock K;Riesselman AJ;Palmedo P;Kang C;Sheridan R;Draizen EJ;Dallago C;Sander C;Marks DS
通讯作者: Marks DS