ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning

ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning
复制标题

DOI:
10.1109/tpami.2021.3095381
复制
发表时间:
2022-10-01
影响因子:
23.6
通讯作者:
Rost, Burkhard
Rost, Burkhard
中科院分区:
计算机科学1区
文献类型:
--
作者:
Elnaggar, Ahmed;Heinzinger, Michael;Rost, Burkhard

文献摘要

被引文献

相似文献

计算生物学和生物信息学从蛋白质序列中提供了巨大的数据金矿,非常适合从自然语言处理(NLP)中提取的语言模型(LM)。这些线性模型以较低的推理成本达到了新的预测前沿。在这里,我们训练了两个自回归模型(Transformer-XL,XLNet)和四个自动编码器模型(BERT,Albert,Electra,T5),这些模型来自UniRef和BFD,包含多达3930亿个氨基酸。蛋白质LM(pLM)在Summit超级计算机上使用5616个GPU和TPU Pod多达1024个核心进行训练。非线性约简揭示了来自未标记数据的原始pLM嵌入捕获了蛋白质序列的一些生物物理特征。我们验证了使用嵌入作为几个后续任务的唯一输入的优势:(1)蛋白质二级结构的每个残基(每个标记)预测(3状态准确度Q3=81%-87%);(2)蛋白质亚细胞位置的每个蛋白质(池化)预测(10状态准确度:Q10=81%)和膜与水溶性(2状态准确度Q2=91%)。对于二级结构,信息量最大的嵌入(ProtT 5)首次超过了最先进的技术,而无需多重序列比对(MSA)或进化信息,从而绕过了昂贵的数据库搜索。总的来说,这些结果意味着pLM学习了一些生命语言的语法。我们所有的模型都可以通过https://github.com/agemagician/ProtTrans获得。
Computational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive models (Transformer-XL, XLNet) and four auto-encoder models (BERT, Albert, Electra, T5) on data from UniRef and BFD containing up to 393 billion amino acids. The protein LMs (pLMs) were trained on the Summit supercomputer using 5616 GPUs and TPU Pod up-to 1024 cores. Dimensionality reduction revealed that the raw pLM-embeddings from unlabeled data captured some biophysical features of protein sequences. We validated the advantage of using the embeddings as exclusive input for several subsequent tasks: (1) a per-residue (per-token) prediction of protein secondary structure (3-state accuracy Q3=81%-87%); (2) per-protein (pooling) predictions of protein sub-cellular location (ten-state accuracy: Q10=81%) and membrane versus water-soluble (2-state accuracy Q2=91%). For secondary structure, the most informative embeddings (ProtT5) for the first time outperformed the state-of-the-art without multiple sequence alignments (MSAs) or evolutionary information thereby bypassing expensive database searches. Taken together, the results implied that pLMs learned some of the grammar of the language of life. All our models are available through https://github.com/agemagician/ProtTrans.