Identification of functional links between genes using phylogenetic profiles

Identification of functional links between genes using phylogenetic profiles
复制标题

DOI:
10.1093/bioinformatics/btg187
复制
发表时间:
2003-08-12
期刊:
影响因子:
5.8
通讯作者:
DeLisi, C
DeLisi, C
中科院分区:
生物学3区
文献类型:
--
作者:
Wu, J;Kasif, S;DeLisi, C

文献摘要

被引文献

相似文献

动机:在不同门中具有相同发生模式的基因倾向于在相同的蛋白质复合物中一起起作用或参与相同的生化途径。然而,要求谱是相同的(i)严重限制了可以通过这种系统发育谱建立的功能链接的数量;(ii)将检测限制到非常强的功能链接,不能捕获不在相同途径中的基因之间的关系,但仍然有助于共同功能;以及(iii)错过类似基因之间的关系。在这里,我们提出并应用一种方法来放松限制,基于概率,即两个配置文件之间的任意程度的相似性会发生的机会,没有生物压力。功能,然后推断在任何所需的水平confidence.Results:我们推导出一个表达式的概率分布的一个给定数量的机会共同出现的一对非同源的直系同源跨一组基因组。该方法适用于2905集群的orthopathic基因(COG)从44个完全测序的微生物基因组代表所有三个领域的生活。结果如下。(1)在51000个注释的途径内基因对中,8935个在0.01的显著性水平上连锁。这是超过30倍以上,在相同的置信水平时,使用相同的配置文件获得的271个通道内对。(2)在540000 interpathway基因对,约65000是在0.01的显着性水平,一些12个标准差超出预期的数量在这个置信水平的机会。我们推测,许多这些链接涉及最近邻路径,并讨论了一些例子。(3)通路间和通路内连锁基因的百分比差异是非常显著的,这与直觉预期一致,即相同通路中的基因通常比不相同通路中的基因受到更大的选择压力。(4)该方法似乎恢复良好的代谢网络。TCA循环说明了这一点,该循环被恢复为31个COG中的30个的高度连接的加权边缘网络。(5)具有共同路径的对的分数是它们的轮廓之间的汉明距离的对称函数。这一发现,即具有接近最大汉明距离的谱之间的功能相关性与具有接近零汉明距离的谱之间的功能相关性一样大,并且具有统计学显著性,如果前一组代表类似基因,则可以合理地解释。
Motivation: Genes with identical patterns of occurrence across the phyla tend to function together in the same protein complexes or participate in the same biochemical pathway. However, the requirement that the profiles be identical (i) severely restricts the number of functional links that can he established by such phylogenetic profiling; (ii) limits detection to very strong functional links, failing to capture relations between genes that are not in the same pathway, but nevertheless subserve a common function and (iii) misses relations between analogous genes. Here we present and apply a method for relaxing the restriction, based on the probability that a given arbitrary degree of similarity between two profiles would occur by chance, with no biological pressure. Function is then inferred at any desired level of confidence.Results: We derive an expression for the probability distribution of a given number of chance co-occurrences of a pair of non-homologous orthologs across a set of genomes. The method is applied to 2905 clusters of orthologous genes (COGs) from 44 fully sequenced microbial genomes representing all three domains of life. Among the results are the following. (1) Of the 51000 annotated intrapathway gene pairs, 8935 are linked at a level of significance of 0.01. This is over 30-fold greater than the 271 intrapathway pairs obtained at the same confidence level when identical profiles are used. (2) Of the 540000 interpathway genes pairs, some 65000 are linked at the 0.01 level of significance, some 12 standard deviations beyond the number expected by chance at this confidence level. We speculate that many of these links involve nearest-neighbor path, and discuss some examples. (3) The difference in the percentage of linked interpathway and intrapathway genes is highly significant, consisten with the intuitive expectation that genes in the same pathway are generally under greater selective pressure than those that are not. (4) The method appears to recover well metabolic networks. This is illustrated by the TCA Cycle which is recovered as a highly connected, weighted edge network of 30 of its 31 COGs. (5) The fraction of pairs having a common pathway is a symmetric function of the Hamming distance between their profiles. This finding, that the functional correlation between profiles with near maximum Hamming distance is as large as between profiles with near zero Hamming distance, and as statistically significant, is plausibly explained if the former group represents analogous genes.