Enhanced protein fold recognition through a novel data integration approach.

Enhanced protein fold recognition through a novel data integration approach.
复制标题

DOI:
10.1186/1471-2105-10-267
复制
发表时间:
2009-08-26
期刊:
影响因子:
3
通讯作者:
Campbell C
Campbell C
中科院分区:
生物学4区
文献类型:
--
作者:
Ying Y;Huang K;Campbell C

文献摘要

参考文献

被引文献

相似文献

蛋白质折叠识别是发现蛋白质三维结构的关键步骤。存在使用物理化学和结构特性的多重判别数据源以及源自局部序列比对的进一步数据源。这就提出了一个问题,即找到最有效的方法来结合这些不同的信息数据源,并探索它们对蛋白质折叠分类的相对意义。核方法已被广泛用于生物数据分析。它们可以将单独的折叠判别特征合并到内核矩阵中,内核矩阵对各自数据源中的样本之间的相似性进行编码。在本文中,我们考虑使用基于内核的方法集成多个数据源的问题。我们提出了一种新的信息论方法的基础上的输出核矩阵和输入核矩阵之间的Kullback-Leibler(KL)分歧,以集成异构数据源。这种方法最吸引人的特性之一是,它可以很容易地科普多类分类和多任务学习的输出核矩阵的适当选择。基于输出和输入核矩阵在KL发散目标中的位置,有两个公式,我们分别称为MKLdiv-dc和MKLdiv-conv。我们建议有效地解决MKLdiv-dc的差异凸(DC)规划方法和MKLdiv-conv的投影梯度下降算法。在蛋白质折叠识别和酵母蛋白质功能预测问题的基准数据集上评估了所提出方法的有效性。我们提出的方法MKLdiv-dc和MKLdiv-conv能够在SCOP PDB-40 D基准数据集上实现最先进的性能,用于蛋白质折叠预测,并为信息数据源的相对重要性提供有用的见解。特别是,MKLdiv-dc进一步将折叠判别准确度提高到75.19%,这比竞争性贝叶斯概率和基于SVM边缘的核学习方法提高了5%以上。此外,我们报告了一个有竞争力的性能上的酵母蛋白质功能预测问题。
Protein fold recognition is a key step in protein three-dimensional (3D) structure discovery. There are multiple fold discriminatory data sources which use physicochemical and structural properties as well as further data sources derived from local sequence alignments. This raises the issue of finding the most efficient method for combining these different informative data sources and exploring their relative significance for protein fold classification. Kernel methods have been extensively used for biological data analysis. They can incorporate separate fold discriminatory features into kernel matrices which encode the similarity between samples in their respective data sources. In this paper we consider the problem of integrating multiple data sources using a kernel-based approach. We propose a novel information-theoretic approach based on a Kullback-Leibler (KL) divergence between the output kernel matrix and the input kernel matrix so as to integrate heterogeneous data sources. One of the most appealing properties of this approach is that it can easily cope with multi-class classification and multi-task learning by an appropriate choice of the output kernel matrix. Based on the position of the output and input kernel matrices in the KL-divergence objective, there are two formulations which we respectively refer to as MKLdiv-dc and MKLdiv-conv. We propose to efficiently solve MKLdiv-dc by a difference of convex (DC) programming method and MKLdiv-conv by a projected gradient descent algorithm. The effectiveness of the proposed approaches is evaluated on a benchmark dataset for protein fold recognition and a yeast protein function prediction problem. Our proposed methods MKLdiv-dc and MKLdiv-conv are able to achieve state-of-the-art performance on the SCOP PDB-40D benchmark dataset for protein fold prediction and provide useful insights into the relative significance of informative data sources. In particular, MKLdiv-dc further improves the fold discrimination accuracy to 75.19% which is a more than 5% improvement over competitive Bayesian probabilistic and SVM margin-based kernel learning methods. Furthermore, we report a competitive performance on the yeast protein function prediction problem.
DOI: 10.1090/s0002-9947-1950-0051437-7
发表时间: 1950-01-01
影响因子: 1.3
作者:
ARONSZAJN, N
通讯作者: ARONSZAJN, N
DOI: 10.1093/bioinformatics/bth294
发表时间: 2004-11-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Lanckriet, GRG;De Bie, T;Noble, WS
通讯作者: Noble, WS
DOI: 10.1073/pnas.92.19.8700
发表时间: 1995-09-12
影响因子: 11.1
作者:
DUBCHAK, I;MUCHNIK, I;KIM, SH
通讯作者: KIM, SH
DOI: 10.1089/106652703322756113
发表时间: 2003-01-01
影响因子: 1.7
作者:
Liao, L;Noble, WS
通讯作者: Noble, WS
DOI: 10.1214/009053606000000722
发表时间: 2006-10-01
影响因子: 4.5
作者:
Lin, Yi;Zhang, Hao Helen
通讯作者: Zhang, Hao Helen