CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models

CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models
复制标题

CATHe:使用蛋白质语言模型的嵌入检测 CATH 超家族的远程同源物

DOI:
10.1101/2022.03.10.483805
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Nallapareddy V
Nallapareddy V
中科院分区:
--
文献类型:
--
作者:
Nallapareddy V

文献摘要

相似文献

MotivationCATH是一个蛋白质结构域分类资源,它利用结构和序列比较的自动化工作流程以及专家手动管理来构建进化和结构关系的分层分类。本研究的目的是开发用于检测由最先进的隐马尔可夫模型(HMM)为基础的方法错过的远程同源物的算法。开发的方法(CATHe)将神经网络与从蛋白质语言模型获得的序列表示相结合。它被评估使用的数据集的远程同源物具有小于20%的序列同一性的任何域在training set.ResultsThe CATHe模型训练1773最大的和50个最大的CATH超家族的准确性分别为85.6 ± 0.4%和98.2 ± 0.3%。作为对CATHe检测由源自CATH结构域的HALTH遗漏的更远同源物的能力的进一步测试,我们使用了由在Pfam中具有注释但在CATH中没有注释的蛋白质结构域组成的数据集。通过使用高度可靠的CATHe预测(预期错误率<0.5%),我们能够为462万个Pfam域提供CATH注释。对于来自智人的这些域的子集,我们通过比较它们相应的AlphaFold 2结构与它们被分配到的CATH超家族的结构,在结构上验证了90.86%的预测。可用性和实施开发模型的代码可在https://github.com/vam-sin/CATHe上获得,本研究中开发的数据集可在https://zenodo.org/record/6327572.Supplementary信息上访问补充数据可在Bioinformaticsonline上获得。
MotivationCATH is a protein domain classification resource that exploits an automated workflow of structure and sequence comparison alongside expert manual curation to construct a hierarchical classification of evolutionary and structural relationships. The aim of this study was to develop algorithms for detecting remote homologues missed by state-of-the-art hidden Markov model (HMM)-based approaches. The method developed (CATHe) combines a neural network with sequence representations obtained from protein language models. It was assessed using a dataset of remote homologues having less than 20% sequence identity to any domain in the training set.ResultsThe CATHe models trained on 1773 largest and 50 largest CATH superfamilies had an accuracy of 85.6 ± 0.4% and 98.2 ± 0.3%, respectively. As a further test of the power of CATHe to detect more remote homologues missed by HMMs derived from CATH domains, we used a dataset consisting of protein domains that had annotations in Pfam, but not in CATH. By using highly reliable CATHe predictions (expected error rate <0.5%), we were able to provide CATH annotations for 4.62 million Pfam domains. For a subset of these domains fromHomo sapiens, we structurally validated 90.86% of the predictions by comparing their corresponding AlphaFold2 structures with structures from the CATH superfamilies to which they were assigned.Availability and implementationThe code for the developed models is available on https://github.com/vam-sin/CATHe, and the datasets developed in this study can be accessed on https://zenodo.org/record/6327572.Supplementary informationSupplementary data are available atBioinformaticsonline.