Large language models generate functional protein sequences across diverse families.
Large language models generate functional protein sequences across diverse families.
复制标题
DOI:
10.1038/s41587-022-01618-2
复制
发表时间:
2023-08
影响因子:
46.9
通讯作者:
Naik, Nikhil
中科院分区:
文献类型:
--
作者:
Madani, Ali;Ben Krause, Ben;Greene, Eric R.;Subramanian, Subu;Mohr, Benjamin P.;Holton, James M.;Olmos, Jose Luis;Xiong, Caiming;Sun, Zachary Z. Z.;Socher, Richard;Fraser, James S.;Naik, Nikhil
Generating artificial protein sequences using artificial intelligence could enable breakthrough solutions for biomedical and environmental challenges. Viewing amino acid sequences as a language, we demonstrate that a deep learning-based language model can generate functional artificial protein sequences across large protein families, akin to generating grammatically and semantically correct natural language sentences on diverse topics. Our protein language model is trained by simply learning to predict the next amino acid for over 280 million protein sequences from thousands of protein families, without explicit biophysical modeling. We experimentally evaluate model-generated artificial proteins fine-tuned to five distinct antibacterial lysozyme families. Artificial proteins show similar activities and catalytic efficiencies as representative natural lysozymes, including hen egg white lysozyme, while maintaining activity with as low as 31.4% identity to any known naturally-evolved protein. The X-ray crystal structure of an enzymatically active artificial protein recapitulates the conserved fold and positioning of active site residues found in natural proteins. We show our language model’s ability to be adapted to different protein families by accurately predicting the functionality of artificial chorismate mutase and malate dehydrogenase proteins. These results indicate that neural language models successfully perform artificial protein generation across protein families and may prove to be a tool to shortcut evolution.
登录
查看更多内容
影响因子:
14.9
作者:
Federhen S
通讯作者:
Federhen S
DOI:
10.1109/tpami.2021.3095381
发表时间:
2022-10-01
影响因子:
23.6
作者:
Elnaggar, Ahmed;Heinzinger, Michael;Rost, Burkhard
通讯作者:
Rost, Burkhard
影响因子:
14.9
作者:
Bairoch, Amos;Bougueleret, Lydie;Zhang, Jian
通讯作者:
Zhang, Jian
影响因子:
2.9
作者:
Balakrishnan, Sivaraman;Kamisetty, Hetunandan;Langmead, Christopher James
通讯作者:
Langmead, Christopher James
DOI:
10.1093/bioinformatics/bty862
发表时间:
2019-05-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
Hopf TA;Green AG;Schubert B;Mersmann S;Schärfe CPI;Ingraham JB;Toth-Petroczy A;Brock K;Riesselman AJ;Palmedo P;Kang C;Sheridan R;Draizen EJ;Dallago C;Sander C;Marks DS
通讯作者:
Marks DS