Sequence-Structure Embeddings via Protein Language Models Improve on Prediction Tasks

Sequence-Structure Embeddings via Protein Language Models Improve on Prediction Tasks
复制标题

DOI:
10.1109/ickg55886.2022.00021
复制
发表时间:
2022-11
期刊:
2022 IEEE International Conference on Knowledge Graph (ICKG)
影响因子:
--
通讯作者:
Anowarul Kabir;Amarda Shehu
Anowarul Kabir;Amarda Shehu
中科院分区:
其他
文献类型:
--
作者:
Anowarul Kabir;Amarda Shehu

文献摘要

相似文献

蛋白质语言模型(PLM)以变压器体系结构及其革命性的语言模型的革命性为基础,现在已经成为一种有力的工具,可以在蛋白质序列数据库中学习大量序列,并将蛋白质序列与功能联系起来。 PLM被证明可以学习有用的任务无关序列表示,可以预测蛋白质二级结构,蛋白质亚细胞定位和蛋白质家族中的进化关系。但是,现有模型是严格培训的,对蛋白质序列进行了严格的培训,并错过了利用和整合异质数据源中存在的信息的机会。在本文中,我们的灵感来自三维/三级蛋白质结构在确定广泛蛋白质特性中的内在作用,我们提出了一个PLM,该PLM会整合并参与蛋白质序列和第三纪结构。特别是,本文认为,学习联合序列结构表示为与功能相关的预测任务提供更好的表示。详细的实验评估表明,这种联合序列结构表示比基于序列的表示功能更强大,在各种指标上产生更好的超家族成员的性能,并在PLM学习的嵌入空间中捕获有趣的关系。
Building on the transformer architecture and its revolutionizing of language models for natural language processing, protein language models (PLMs) are now emerging as a powerful tool for learning over large numbers of sequences in protein sequence databases and linking protein sequence to function. PLMs are shown to learn useful, task-agnostic sequence representations that allow predicting protein secondary structure, protein subcellular localization, and evolutionary relationships within protein families. However, existing models are strictly trained over protein sequences and miss an opportunity to leverage and integrate the information present in heterogeneous data sources. In this paper, inspired by the intrinsic role of three-dimensional/tertiary protein structure in determining a broad range of protein properties, we propose a PLM that integrates and attends to both protein sequence and tertiary structure. In particular, this paper posits that learning joint sequence-structure representations yields better representations for function-related prediction tasks. A detailed experimental evaluation shows that such joint sequence-structure representations are more powerful than sequence-based representations, yield better performance on superfamily membership across various metrics, and capture interesting relationships in the PLM-learned embedding space.