Tree visualizations of protein sequence embedding space enable improved functional clustering of diverse protein superfamilies.

Tree visualizations of protein sequence embedding space enable improved functional clustering of diverse protein superfamilies.
复制标题

DOI:
10.1093/bib/bbac619
复制
发表时间:
2023-01-19
影响因子:
9.5
通讯作者:
--
中科院分区:
生物学2区
文献类型:
--
作者:

文献摘要

参考文献

相似文献

蛋白质语言模型经过数百万个生物观察序列的训练,生成蛋白质序列的特征丰富的数字表示。尽管蛋白质语言模型仅在初级序列上进行训练,但这些表示(称为序列嵌入)可以推断结构功能特性。虽然序列嵌入已应用于结构和功能预测等任务,但由于缺乏推导、量化和评估蛋白质序列嵌入之间关系的研究,免对齐序列分类的应用受到阻碍。在这里,我们使用从蛋白质语言模型派生的序列嵌入来开发用于蛋白质家族分类的工作流程和可视化方法。流形可视化方法的基准表明,与流行的降维技术(如 t-SNE 和 UMAP)相比,邻接(NJ)嵌入树在捕获全局结构方面非常有效,同时在捕获局部结构方面实现了相似的性能。通过使用变分自动编码器(VAE)对嵌入进行重采样来评估树上分层聚类的统计显着性。我们展示了我们的方法在两个经过充分研究的酶超家族(磷酸酶和蛋白激酶)的分类中的应用。我们基于嵌入的分类与之前发布的基于序列比对的分类保持一致并扩展。我们还为 S-腺苷-L-甲硫氨酸 (SAM) 酶超家族提出了一种新的层次分类,该家族很难使用传统的基于比对的方法进行分类。除了序列分类中的应用之外,我们的结果进一步表明 NJ 树是一种有前途的高维数据集可视化通用方法。
Protein language models, trained on millions of biologically observed sequences, generate feature-rich numerical representations of protein sequences. These representations, called sequence embeddings, can infer structure-functional properties, despite protein language models being trained on primary sequence alone. While sequence embeddings have been applied toward tasks such as structure and function prediction, applications toward alignment-free sequence classification have been hindered by the lack of studies to derive, quantify and evaluate relationships between protein sequence embeddings. Here, we develop workflows and visualization methods for the classification of protein families using sequence embedding derived from protein language models. A benchmark of manifold visualization methods reveals that Neighbor Joining (NJ) embedding trees are highly effective in capturing global structure while achieving similar performance in capturing local structure compared with popular dimensionality reduction techniques such as t-SNE and UMAP. The statistical significance of hierarchical clusters on a tree is evaluated by resampling embeddings using a variational autoencoder (VAE). We demonstrate the application of our methods in the classification of two well-studied enzyme superfamilies, phosphatases and protein kinases. Our embedding-based classifications remain consistent with and extend upon previously published sequence alignment-based classifications. We also propose a new hierarchical classification for the S-Adenosyl-L-Methionine (SAM) enzyme superfamily which has been difficult to classify using traditional alignment-based approaches. Beyond applications in sequence classification, our results further suggest NJ trees are a promising general method for visualizing high-dimensional data sets.
DOI: 10.1038/nchembio.1426
发表时间: 2014-02
影响因子: 14.8
作者:
Dowling, Daniel P.;Bruender, Nathan A.;Young, Anthony P.;McCarty, Reid M.;Bandarian, Vahe;Drennan, Catherine L.
通讯作者: Drennan, Catherine L.
DOI: 10.1126/scisignal.aag1796
发表时间: 2017-04-11
期刊: SCIENCE SIGNALING
影响因子: 7.3
作者:
Chen, Mark J.;Dixon, Jack E.;Manning, Gerard
通讯作者: Manning, Gerard
DOI: 10.1109/tpami.2021.3095381
发表时间: 2022-10-01
影响因子: 23.6
作者:
Elnaggar, Ahmed;Heinzinger, Michael;Rost, Burkhard
通讯作者: Rost, Burkhard
DOI: 10.1016/0304-4165(85)90034-0
发表时间: 1985-01-01
期刊: BIOCHIMICA ET BIOPHYSICA ACTA
影响因子: --
作者:
CHAKRABARTTY, A;STINSON, RA
通讯作者: STINSON, RA
DOI: 10.1038/s41586-020-2762-2
发表时间: 2021-01
期刊: Nature
影响因子: 64.8
作者:
Bernheim A;Millman A;Ofir G;Meitav G;Avraham C;Shomar H;Rosenberg MM;Tal N;Melamed S;Amitai G;Sorek R
通讯作者: Sorek R