CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models.

CATHe: detection of remote homologues for CATH superfamilies using embeddings from protein language models.
复制标题

CATHe:使用蛋白质语言模型的嵌入检测 CATH 超家族的远程同源物。

DOI:
10.1093/bioinformatics/btad029
复制
发表时间:
2023-01-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

CATH是一个蛋白质结构域分类资源库,它利用结构与序列比较的自动化流程,结合专家人工审编,构建进化与结构关系的层级分类。本研究旨在开发算法,以检测当前基于隐马尔可夫模型(HMM)的先进方法所遗漏的远源同源物。所开发的方法(CATHe)将神经网络与从蛋白质语言模型获得的序列表征相结合。该方法使用一个远源同源物数据集进行评估,这些远源同源物与训练集中任何结构域的序列同一性均低于20%。 在1773个最大和50个最大的CATH超家族上训练的CATHe模型,准确率分别为85.6±0.4%和98.2±0.3%。为进一步测试CATHe检测CATH结构域衍生的HMM所遗漏的更多远源同源物的能力,我们使用了一个蛋白质结构域数据集,这些结构域在Pfam中有注释,但在CATH中没有。通过使用高度可靠的CATHe预测(预期错误率<0.5%),我们能够为462万个Pfam结构域提供CATH注释。对于来自智人的这些结构域的一个子集,我们通过将其相应的AlphaFold2结构与它们所归属的CATH超家族的结构进行比较,对90.86%的预测进行了结构验证。 所开发模型的代码可在https://github.com/vam-sin/CATHe获取,本研究中开发的数据集可在https://zenodo.org/record/6327572访问。 补充数据可在《生物信息学》在线版获取。
CATH is a protein domain classification resource that exploits an automated workflow of structure and sequence comparison alongside expert manual curation to construct a hierarchical classification of evolutionary and structural relationships. The aim of this study was to develop algorithms for detecting remote homologues missed by state-of-the-art hidden Markov model (HMM)-based approaches. The method developed (CATHe) combines a neural network with sequence representations obtained from protein language models. It was assessed using a dataset of remote homologues having less than 20% sequence identity to any domain in the training set. The CATHe models trained on 1773 largest and 50 largest CATH superfamilies had an accuracy of 85.6 ± 0.4% and 98.2 ± 0.3%, respectively. As a further test of the power of CATHe to detect more remote homologues missed by HMMs derived from CATH domains, we used a dataset consisting of protein domains that had annotations in Pfam, but not in CATH. By using highly reliable CATHe predictions (expected error rate <0.5%), we were able to provide CATH annotations for 4.62 million Pfam domains. For a subset of these domains from Homo sapiens, we structurally validated 90.86% of the predictions by comparing their corresponding AlphaFold2 structures with structures from the CATH superfamilies to which they were assigned. The code for the developed models is available on https://github.com/vam-sin/CATHe, and the datasets developed in this study can be accessed on https://zenodo.org/record/6327572. Supplementary data are available at Bioinformatics online.
塞思预测蛋白质嵌入的残基障碍的细微差别。
DOI: 10.3389/fbinf.2022.1019597
发表时间: 2022
期刊: FRONTIERS IN BIOINFORMATICS
影响因子: --
作者:
Ilzhofer, Dagmar;Heinzinger, Michael;Rost, Burkhard
通讯作者: Rost, Burkhard
DOI: 10.1038/s41586-021-03819-2
发表时间: 2021-08
期刊: Nature
影响因子: 64.8
作者:
Jumper J;Evans R;Pritzel A;Green T;Figurnov M;Ronneberger O;Tunyasuvunakool K;Bates R;Žídek A;Potapenko A;Bridgland A;Meyer C;Kohl SAA;Ballard AJ;Cowie A;Romera-Paredes B;Nikolov S;Jain R;Adler J;Back T;Petersen S;Reiman D;Clancy E;Zielinski M;Steinegger M;Pacholska M;Berghammer T;Bodenstein S;Silver D;Vinyals O;Senior AW;Kavukcuoglu K;Kohli P;Hassabis D
通讯作者: Hassabis D
DOI: 10.1093/nar/gkaa913
发表时间: 2021-01-08
影响因子: 14.9
作者:
Mistry J;Chuguransky S;Williams L;Qureshi M;Salazar GA;Sonnhammer ELL;Tosatto SCE;Paladin L;Raj S;Richardson LJ;Finn RD;Bateman A
通讯作者: Bateman A
DOI: 10.1093/nar/gky949
发表时间: 2019-01-08
影响因子: 14.9
作者:
wwPDB consortium
通讯作者: wwPDB consortium
DOI: 10.1093/nar/gky1100
发表时间: 2019-01-08
影响因子: 14.9
作者:
Mitchell AL;Attwood TK;Babbitt PC;Blum M;Bork P;Bridge A;Brown SD;Chang HY;El-Gebali S;Fraser MI;Gough J;Haft DR;Huang H;Letunic I;Lopez R;Luciani A;Madeira F;Marchler-Bauer A;Mi H;Natale DA;Necci M;Nuka G;Orengo C;Pandurangan AP;Paysan-Lafosse T;Pesseat S;Potter SC;Qureshi MA;Rawlings ND;Redaschi N;Richardson LJ;Rivoire C;Salazar GA;Sangrador-Vegas A;Sigrist CJA;Sillitoe I;Sutton GG;Thanki N;Thomas PD;Tosatto SCE;Yong SY;Finn RD
通讯作者: Finn RD