Clustering FunFams using sequence embeddings improves EC purity.

Clustering FunFams using sequence embeddings improves EC purity.
复制标题

DOI:
10.1093/bioinformatics/btab371
复制
发表时间:
2021-10-25
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
Rost B
Rost B
中科院分区:
其他
文献类型:
--
作者:
Littmann M;Bordin N;Heinzinger M;Schütze K;Dallago C;Orengo C;Rost B

文献摘要

参考文献

被引文献

相似文献

将蛋白质分类为功能家族可以提高我们对蛋白质功能的理解,并可以在一个家族内转移注释。为此,功能家族必须是“纯粹的”,即只包含具有相同功能的蛋白质。功能家族(FunFams)将CATH超家族中的蛋白质聚集成共享功能的蛋白质群。11%的FunFams(203 639个中的22 830个)包含EC注释,其中7%(22 830个中的1526个)具有不一致的功能注释。我们提出了一种方法,通过嵌入编码它们的序列,进一步将FunFams聚类成功能更一致的子家族。这些嵌入源自语言模型,通过预测序列中缺失的氨基酸而获得知识(ProtBERT),并已进一步优化以区分属于相同或不同CATH超家族的蛋白质(PB-Tucker)。使用嵌入和DBSCAN之间的距离对FunFam进行聚类并识别异常值,与随机聚类相比,每个FunFam的纯聚类数量增加了一倍。我们的方法不仅局限于FunFams,而且在仅使用序列相似性创建的家庭中也取得了成功。作为EC注释的补充,我们在绑定注释中观察到类似的结果。因此,我们期望在功能的其他方面也能提高纯度。我们的研究结果可以帮助创建FunFams;得到的具有改进功能一致性的聚类允许对注释进行更可靠的推断。我们期望这种方法同样成功的任何其他组的蛋白质的表型。代码和嵌入可通过GitHub: https://github.com/Rostlab/FunFamsClustering。补充数据可在生物信息学网站获得。
Classifying proteins into functional families can improve our understanding of protein function and can allow transferring annotations within one family. For this, functional families need to be ‘pure’, i.e., contain only proteins with identical function. Functional Families (FunFams) cluster proteins within CATH superfamilies into such groups of proteins sharing function. 11% of all FunFams (22 830 of 203 639) contain EC annotations and of those, 7% (1526 of 22 830) have inconsistent functional annotations. We propose an approach to further cluster FunFams into functionally more consistent sub-families by encoding their sequences through embeddings. These embeddings originate from language models transferring knowledge gained from predicting missing amino acids in a sequence (ProtBERT) and have been further optimized to distinguish between proteins belonging to the same or a different CATH superfamily (PB-Tucker). Using distances between embeddings and DBSCAN to cluster FunFams and identify outliers, doubled the number of pure clusters per FunFam compared to random clustering. Our approach was not limited to FunFams but also succeeded on families created using sequence similarity alone. Complementing EC annotations, we observed similar results for binding annotations. Thus, we expect an increased purity also for other aspects of function. Our results can help generating FunFams; the resulting clusters with improved functional consistency allow more reliable inference of annotations. We expect this approach to succeed equally for any other grouping of proteins by their phenotypes. Code and embeddings are available via GitHub: https://github.com/Rostlab/FunFamsClustering. Supplementary data are available at Bioinformatics online.
DOI: 10.1093/bioinformatics/btn214
发表时间: 2008-07-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者:
Capra JA;Singh M
通讯作者: Singh M
DOI: 10.1109/tpami.2021.3095381
发表时间: 2022-10-01
影响因子: 23.6
作者:
Elnaggar, Ahmed;Heinzinger, Michael;Rost, Burkhard
通讯作者: Rost, Burkhard
DOI: 10.1038/s41592-019-0598-1
发表时间: 2019-12-01
期刊: NATURE METHODS
影响因子: 48
作者:
Alley, Ethan C.;Khimulya, Grigory;Church, George M.
通讯作者: Church, George M.
DOI: 10.1371/journal.pone.0141287
发表时间: 2015
期刊: PloS one
影响因子: 3.7
作者:
Asgari E;Mofrad MR
通讯作者: Mofrad MR
DOI: 10.1038/s41598-020-80786-0
发表时间: 2021-01-13
期刊: Scientific reports
影响因子: 4.6
作者:
Littmann M;Heinzinger M;Dallago C;Olenyi T;Rost B
通讯作者: Rost B