Learning sequence, structure, and function representations of proteins with language models.

Learning sequence, structure, and function representations of proteins with language models.
复制标题

使用语言模型学习蛋白质的序列、结构和功能表示。

DOI:
10.1101/2023.11.26.568742
复制
发表时间:
2023
期刊:
bioRxiv : the preprint server for biology
影响因子:
--
通讯作者:
Cho,Kyunghyun
Cho,Kyunghyun
中科院分区:
--
文献类型:
--
作者:
Hamamsy,Tymor;Barot,Meet;Morton,JamesT;Steinegger,Martin;Bonneau,Richard;Cho,Kyunghyun

文献摘要

相似文献

最终产生现存观察到的蛋白质多样性的序列-结构-功能关系是复杂的,因为蛋白质弥合了几乎所有细胞过程中涉及的多个信息和物理尺度之间的差距。现有的蛋白质注释数据库如UniProt的一个局限性是只有不到1%的蛋白质具有经过实验验证的功能,需要计算方法来填补缺失的信息。在这里,我们证明了基于蛋白质语言模型的多方面框架可以学习氨基酸序列的序列-结构-功能表示,并可以为敏感的序列-结构-功能感知的蛋白质序列搜索和注释提供基础。基于这个模型,我们介绍了一个多方面的蛋白质信息检索系统,蛋白质-VEC,涵盖了序列、结构和功能方面,可以在生命树的尺度上进行蛋白质的计算注释和功能预测。
The sequence-structure-function relationships that ultimately generate the diversity of extant observed proteins is complex, as proteins bridge the gap between multiple informational and physical scales involved in nearly all cellular processes. One limitation of existing protein annotation databases such as UniProt is that less than 1% of proteins have experimentally verified functions, and computational methods are needed to fill in the missing information. Here, we demonstrate that a multi-aspect framework based on protein language models can learn sequence-structure-function representations of amino acid sequences, and can provide the foundation for sensitive sequence-structure-function aware protein sequence search and annotation. Based on this model, we introduce a multi-aspect information retrieval system for proteins, Protein-Vec, covering sequence, structure, and function aspects, that enables computational protein annotation and function prediction at tree-of-life scales.