Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detection

Combining evolutionary information extracted from frequency profiles with sequence-based kernels for protein remote homology detection
复制标题

将从频率分布中提取的进化信息与基于序列的内核相结合以进行蛋白质远程同源性检测

DOI:
10.1093/bioinformatics/btt709
复制
发表时间:
2014-02-15
期刊:
影响因子:
5.8
通讯作者:
Chou, Kuo-Chen
Chou, Kuo-Chen
中科院分区:
生物学3区
文献类型:
--
作者:
Liu, Bin;Zhang, Deyuan;Chou, Kuo-Chen

文献摘要

被引文献

相似文献

抽象动机:由于其在基础研究(如分子进化和蛋白质属性预测)和实际应用(如药物开发所需的蛋白质三维结构的实时建模)中的重要性,蛋白质远程同源性检测引起了人们极大的兴趣。值得注意的是,基于概况的方法在这方面很有前途,潜力很大。为了进一步提高蛋白质远程同源性检测的效率,如何找到一种最佳的方法将进化信息提取到图谱中是一个关键步骤。结果如下:在这里,我们提出了一种新的方法,所谓的基于配置文件的蛋白质表示,通过频率配置文件提取的进化信息。后者可以从PSI-BLAST产生的多重序列比对计算。三个性能最好的基于序列的内核(SVM-Ngram,SVM-pairwise和SVM-LA)与基于轮廓的蛋白质表示相结合。对包含54个家族和23个超家族的SCOP基准数据集进行了各种测试。实验结果表明,该方法具有良好的应用前景,能明显提高三种核函数的性能。此外,我们的方法也可以为研究不同家族蛋白质的特征提供有用的见解。我们注意到,目前的方法可以很容易地与现有的基于序列的方法相结合,以提高其性能。可用性和实施:为了方便用户,还在http://bioinformatics.hitsz.edu.cn/main/jibinliu/remote/提供了生成基于谱的蛋白质和多核学习的源代码,联系方式:bliu@insun.hit.edu.cn或bliu@gordonlifescience.org补充信息:补充数据可在Bioinformatics在线获得。
Abstract Motivation: Owing to its importance in both basic research (such as molecular evolution and protein attribute prediction) and practical application (such as timely modeling the 3D structures of proteins targeted for drug development), protein remote homology detection has attracted a great deal of interest. It is intriguing to note that the profile-based approach is promising and holds high potential in this regard. To further improve protein remote homology detection, a key step is how to find an optimal means to extract the evolutionary information into the profiles. Results: Here, we propose a novel approach, the so-called profile-based protein representation, to extract the evolutionary information via the frequency profiles. The latter can be calculated from the multiple sequence alignments generated by PSI-BLAST. Three top performing sequence-based kernels (SVM-Ngram, SVM-pairwise and SVM-LA) were combined with the profile-based protein representation. Various tests were conducted on a SCOP benchmark dataset that contains 54 families and 23 superfamilies. The results showed that the new approach is promising, and can obviously improve the performance of the three kernels. Furthermore, our approach can also provide useful insights for studying the features of proteins in various families. It has not escaped our notice that the current approach can be easily combined with the existing sequence-based methods so as to improve their performance as well. Availability and implementation: For users’ convenience, the source code of generating the profile-based proteins and the multiple kernel learning was also provided at http://bioinformatics.hitsz.edu.cn/main/∼binliu/remote/ Contact: bliu@insun.hit.edu.cn or bliu@gordonlifescience.org Supplementary information: Supplementary data are available at Bioinformatics online.