Discovering functional linkages and uncharacterized cellular pathways using phylogenetic profile comparisons: a comprehensive assessment.

Discovering functional linkages and uncharacterized cellular pathways using phylogenetic profile comparisons: a comprehensive assessment.
复制标题

DOI:
10.1186/1471-2105-8-173
复制
发表时间:
2007-05-23
期刊:
影响因子:
3
通讯作者:
Aravind L
Aravind L
中科院分区:
生物学4区
文献类型:
--
作者:
Jothi R;Przytycka TM;Aravind L

文献摘要

参考文献

被引文献

相似文献

一个广泛使用的方法来发现蛋白质之间的功能和物理相互作用涉及系统发育谱比较(PPC)。在这里,具有相似特征的蛋白质被推断为在功能上相关,假设参与相同代谢途径或细胞系统的蛋白质可能在进化过程中共同遗传。我们用E.具有16种不同的精心组成的参考基因组的大肠杆菌和酵母蛋白质表明,仅原核生物中蛋白质的系统发育模式就足以做出相当准确的功能连锁预测。在将少量真核生物添加到参考组中时观察到性能的轻微改善,但是随着真核生物数量的增加观察到性能的显著下降。将大多数寄生虫、病原体或脊椎动物基因组以及相同物种的多个菌株纳入参考集中不一定有助于提高灵敏度或准确度。有趣的是,我们还发现,个别途径的进化历史有显着影响的PPC方法的性能相对于一个特定的参考集。例如,为了准确预测碳水化合物或脂质代谢中的功能链接,与由所有三个超级王国的基因组组成的基因组相比,仅由原核(或细菌)基因组组成的参考组表现最好;这与预测翻译中的功能链接相反,由原核(或细菌)基因组组成的参考组表现最差。我们还证明了广泛使用的随机零模型来量化轮廓相似性的统计显著性是不完整的,这可能会导致假阳性的数量增加。与以前的建议相反,它不仅仅是基因组的数量,而是在参考集中仔细选择信息基因组,影响PPC方法的预测准确性。我们注意到PPC方法的预测能力,特别是在真核生物中,受到主要内共生和随后细菌贡献的严重影响。寄生单细胞真核生物和脊椎动物的过度代表性另外使得真核生物在参考组中不太有用。由来自所有三个超级王国的高度非冗余的基因组集合组成的参考集合在显示相当大的垂直遗传和强保守性(例如翻译装置)的途径中表现得更好,而仅由原核基因组组成的参考集合在更可变的途径如碳水化合物代谢中表现得更好。PPC方法在各种途径上的差异表现,以及功能和轮廓相似性之间的弱正相关性表明,在解释从使用单个参考集的全基因组大规模轮廓比较推断的功能联系时应谨慎。
A widely-used approach for discovering functional and physical interactions among proteins involves phylogenetic profile comparisons (PPCs). Here, proteins with similar profiles are inferred to be functionally related under the assumption that proteins involved in the same metabolic pathway or cellular system are likely to have been co-inherited during evolution. Our experimentation with E. coli and yeast proteins with 16 different carefully composed reference sets of genomes revealed that the phyletic patterns of proteins in prokaryotes alone could be adequate enough to make reasonably accurate functional linkage predictions. A slight improvement in performance is observed on adding few eukaryotes into the reference set, but a noticeable drop-off in performance is observed with increased number of eukaryotes. Inclusion of most parasitic, pathogenic or vertebrate genomes and multiple strains of the same species into the reference set do not necessarily contribute to an improved sensitivity or accuracy. Interestingly, we also found that evolutionary histories of individual pathways have a significant affect on the performance of the PPC approach with respect to a particular reference set. For example, to accurately predict functional links in carbohydrate or lipid metabolism, a reference set solely composed of prokaryotic (or bacterial) genomes performed among the best compared to one composed of genomes from all three super-kingdoms; this is in contrast to predicting functional links in translation for which a reference set composed of prokaryotic (or bacterial) genomes performed the worst. We also demonstrate that the widely used random null model to quantify the statistical significance of profile similarity is incomplete, which could result in an increased number of false-positives. Contrary to previous proposals, it is not merely the number of genomes but a careful selection of informative genomes in the reference set that influences the prediction accuracy of the PPC approach. We note that the predictive power of the PPC approach, especially in eukaryotes, is heavily influenced by the primary endosymbiosis and subsequent bacterial contributions. The over-representation of parasitic unicellular eukaryotes and vertebrates additionally make eukaryotes less useful in the reference sets. Reference sets composed of highly non-redundant set of genomes from all three super-kingdoms fare better with pathways showing considerable vertical inheritance and strong conservation (e.g. translation apparatus), while reference sets solely composed of prokaryotic genomes fare better for more variable pathways like carbohydrate metabolism. Differential performance of the PPC approach on various pathways, and a weak positive correlation between functional and profile similarities suggest that caution should be exercised while interpreting functional linkages inferred from genome-wide large-scale profile comparisons using a single reference set.
DOI: 10.1101/gr.4336406
发表时间: 2006-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Campillos, M;von Mering, C;Bork, P
通讯作者: Bork, P
DOI: 10.1073/pnas.0402591101
发表时间: 2004-06-15
影响因子: 11.1
作者:
Fraser, HB;Hirsh, AE;Eisen, MB
通讯作者: Eisen, MB
DOI: 10.1093/bioinformatics/bti313
发表时间: 2005-05-15
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Date, SV;Marcotte, EM
通讯作者: Marcotte, EM
DOI: 10.1038/415141a
发表时间: 2002-01-10
期刊: NATURE
影响因子: 64.8
作者:
Gavin, AC;Bösche, M;Superti-Furga, G
通讯作者: Superti-Furga, G
DOI: 10.1016/s0968-0004(98)01274-2
发表时间: 1998-09-01
影响因子: 13.8
作者:
Dandekar, T;Snel, B;Bork, P
通讯作者: Bork, P