Phylogenetic tree information aids supervised learning for predicting protein-protein interaction based on distance matrices

Phylogenetic tree information aids supervised learning for predicting protein-protein interaction based on distance matrices
复制标题

DOI:
10.1186/1471-2105-8-6
复制
发表时间:
2007-01-09
期刊:
影响因子:
3
通讯作者:
Liao, Li
Liao, Li
中科院分区:
生物学4区
文献类型:
--
作者:
Craig, Roger A.;Liao, Li

文献摘要

被引文献

相似文献

背景:蛋白质-蛋白质相互作用对细胞功能至关重要。最近开发的预测蛋白质-蛋白质相互作用的计算方法利用相互作用伙伴的共同进化信息,例如。例如,在一个实施例中,距离矩阵之间的相关性,其中每个矩阵存储蛋白质与其来自一组参考genomes.Results的直系同源物之间的成对距离:我们提出了一种新的,简单的方法来解释一些矩阵内的相关性,以提高预测精度。具体地,参考基因组的系统发育物种树被用作用于直链淀粉蛋白的分层聚类的指导树。这些聚类之间的距离,从原始的成对距离矩阵使用邻居连接算法,形成中间距离矩阵,然后将其转换并连接成一个超级系统发育向量。支持向量机在蛋白质对上进行训练和测试,表示为超级系统发育向量,其相互作用是已知的。在交叉验证实验中,以ROC得分衡量的性能显示了我们方法的显着改进(ROC评分0.8446)优于Pearson相关性(0.6587).结论:我们已经表明,系统发育树可以作为一个指导,以提取矩阵内的相关性的距离矩阵的正交蛋白质,其中这些相关性表示为祖先正向同源蛋白的中间距离矩阵。无监督和监督学习范式都受益于这些中间距离矩阵的明确包含,特别是在后一种情况下,这在预测蛋白质-蛋白质相互作用时提供了灵敏度和特异性之间的更好平衡。
Background: Protein-protein interactions are critical for cellular functions. Recently developed computational approaches for predicting protein-protein interactions utilize co-evolutionary information of the interacting partners, e. g., correlations between distance matrices, where each matrix stores the pairwise distances between a protein and its orthologs from a group of reference genomes.Results: We proposed a novel, simple method to account for some of the intra-matrix correlations in improving the prediction accuracy. Specifically, the phylogenetic species tree of the reference genomes is used as a guide tree for hierarchical clustering of the orthologous proteins. The distances between these clusters, derived from the original pairwise distance matrix using the Neighbor Joining algorithm, form intermediate distance matrices, which are then transformed and concatenated into a super phylogenetic vector. A support vector machine is trained and tested on pairs of proteins, represented as super phylogenetic vectors, whose interactions are known. The performance, measured as ROC score in cross validation experiments, shows significant improvement of our method ( ROC score 0.8446) over that of using Pearson correlations ( 0.6587).Conclusion: We have shown that the phylogenetic tree can be used as a guide to extract intramatrix correlations in the distance matrices of orthologous proteins, where these correlations are represented as intermediate distance matrices of the ancestral orthologous proteins. Both the unsupervised and supervised learning paradigms benefit from the explicit inclusion of these intermediate distance matrices, and particularly so in the latter case, which offers a better balance between sensitivity and specificity in the prediction of protein-protein interactions.