FUpred: detecting protein domains through deep-learning-based contact map prediction

FUpred: detecting protein domains through deep-learning-based contact map prediction
复制标题

DOI:
10.1093/bioinformatics/btaa217
复制
发表时间:
2020-06-15
期刊:
影响因子:
5.8
通讯作者:
Zhang, Yang
Zhang, Yang
中科院分区:
生物学3区
文献类型:
--
作者:
Zheng, Wei;Zhou, Xiaogen;Zhang, Yang

文献摘要

被引文献

相似文献

动机:蛋白质结构域是可以独立折叠和起作用的亚基。因此,正确的结构域边界分配是准确分析蛋白质结构和功能的关键一步。然而,目前还没有一种有效的算法可以从序列中进行精确的区域预测。对于具有不连续结构域的蛋白质来说,这个问题尤其具有挑战性,这些结构域由沿序列分离的结构域片段组成。结果:我们开发了一种新的算法FUpred,该算法利用深度残差神经网络与协同进化精度矩阵创建的接触图来预测蛋白质结构域边界。该算法的核心思想是通过最大化域内接触的数量来检索域边界位置,同时最小化域间接触的数量。FUpred在包含2549种蛋白质的大规模数据集上进行了测试,并生成了正确的单域和多域分类,其马修相关系数为0.799,比最佳的基于机器学习(或线程)的方法高19.1%(或5.3%)。对于结构域不连续的蛋白,FUpred的结构域边界检测和归一化结构域重叠得分分别为0.788和0.521,分别比最佳对照方法提高17.3%和23.8%。研究结果为从序列中准确检测结构域组成提供了新的途径,特别是对于不连续的多结构域蛋白质。
Motivation: Protein domains are subunits that can fold and function independently. Correct domain boundary assignment is thus a critical step toward accurate protein structure and function analyses. There is, however, no efficient algorithm available for accurate domain prediction from sequence. The problem is particularly challenging for proteins with discontinuous domains, which consist of domain segments that are separated along the sequence.Results: We developed a new algorithm, FUpred, which predicts protein domain boundaries utilizing contact maps created by deep residual neural networks coupled with coevolutionary precision matrices. The core idea of the algorithm is to retrieve domain boundary locations by maximizing the number of intra-domain contacts, while minimizing the number of inter-domain contacts from the contact maps. FUpred was tested on a large-scale dataset consisting of 2549 proteins and generated correct single- and multi-domain classifications with a Matthew's correlation coefficient of 0.799, which was 19.1% (or 5.3%) higher than the best machine learning (or threading)-based method. For proteins with discontinuous domains, the domain boundary detection and normalized domain overlapping scores of FUpred were 0.788 and 0.521, respectively, which were 17.3% and 23.8% higher than the best control method. The results demonstrate a new avenue to accurately detect domain composition from sequence alone, especially for discontinuous, multi-domain proteins.