Assessing performance of orthology detection strategies applied to eukaryotic genomes.

Assessing performance of orthology detection strategies applied to eukaryotic genomes.
复制标题

DOI:
10.1371/journal.pone.0000383
复制
发表时间:
2007-04-18
期刊:
影响因子:
3.7
通讯作者:
Roos DS
Roos DS
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Chen F;Mackey AJ;Vermunt JK;Roos DS

文献摘要

参考文献

被引文献

相似文献

同源性检测对于精确的功能注释至关重要,并已被广泛用于促进比较和进化基因组学的研究。虽然现在有各种各样的方法,但由于缺乏基因组规模的“金标准”直系数据集,还没有全面的性能分析。即使在没有这种数据集的情况下,对替代方法的结果进行比较也包含有用的信息,因为一致可以增强信心,不一致则表明可能存在错误。潜在类别分析(LCA)是一种统计技术,可以利用这些信息来合理地推断灵敏度和特异性,并在这里应用于评估真核生物数据集上各种同源检测方法的性能。总的来说,我们观察到同源性检测的灵敏度和特异性之间的权衡,基于BLAST的方法具有高灵敏度的特征,基于树的方法具有高特异性。两种算法表现出最佳的总体平衡,灵敏度和特异性均>80%:INPARANOID识别两个物种的同源物,而OrthoMCL聚类来自多个物种的同源物。在允许跨越多个基因组的直系同源物组聚类的方法中,(自动化)OrthoMCL算法在蛋白质功能和结构域架构方面表现出比(手动策划的)KOG数据库更好的组内一致性,以及同源物聚类算法TribeMCL。通过使用LCA,我们还能够全面评估各种策略之间的相似性和统计依赖性,并评估参数设置对性能的影响。总之,我们提出了一个全面的评估,一组不同的真核生物基因组的同源性检测,从而为不同的应用程序的方法选择,调整和开发提供见解和指导。许多生物学问题已经解决了多个测试产生的二进制(是/否)的结果,但没有明确的定义的真理,使LCA计算生物学的一个有吸引力的方法。
Orthology detection is critically important for accurate functional annotation, and has been widely used to facilitate studies on comparative and evolutionary genomics. Although various methods are now available, there has been no comprehensive analysis of performance, due to the lack of a genomic-scale ‘gold standard’ orthology dataset. Even in the absence of such datasets, the comparison of results from alternative methodologies contains useful information, as agreement enhances confidence and disagreement indicates possible errors. Latent Class Analysis (LCA) is a statistical technique that can exploit this information to reasonably infer sensitivities and specificities, and is applied here to evaluate the performance of various orthology detection methods on a eukaryotic dataset. Overall, we observe a trade-off between sensitivity and specificity in orthology detection, with BLAST-based methods characterized by high sensitivity, and tree-based methods by high specificity. Two algorithms exhibit the best overall balance, with both sensitivity and specificity>80%: INPARANOID identifies orthologs across two species while OrthoMCL clusters orthologs from multiple species. Among methods that permit clustering of ortholog groups spanning multiple genomes, the (automated) OrthoMCL algorithm exhibits better within-group consistency with respect to protein function and domain architecture than the (manually curated) KOG database, and the homolog clustering algorithm TribeMCL as well. By way of using LCA, we are also able to comprehensively assess similarities and statistical dependence between various strategies, and evaluate the effects of parameter settings on performance. In summary, we present a comprehensive evaluation of orthology detection on a divergent set of eukaryotic genomes, thus providing insights and guides for method selection, tuning and development for different applications. Many biological questions have been addressed by multiple tests yielding binary (yes/no) outcomes but no clear definition of truth, making LCA an attractive approach for computational biology.
DOI: 10.1101/gr.212002
发表时间: 2002-03-01
期刊: GENOME RESEARCH
影响因子: 7
作者:
Lee, Y;Sultana, R;Quackenbush, J
通讯作者: Quackenbush, J
DOI: 10.2307/2412448
发表时间: 1970-01-01
期刊: SYSTEMATIC ZOOLOGY
影响因子: --
作者:
FITCH, WM
通讯作者: FITCH, WM
DOI: 10.1186/gb-2004-5-2-r7
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Koonin EV;Fedorova ND;Jackson JD;Jacobs AR;Krylov DM;Makarova KS;Mazumder R;Mekhedov SL;Nikolskaya AN;Rao BS;Rogozin IB;Smirnov S;Sorokin AV;Sverdlov AV;Vasudevan S;Wolf YI;Yin JJ;Natale DA
通讯作者: Natale DA
DOI: 10.1126/science.278.5338.609
发表时间: 1997-10-24
期刊: SCIENCE
影响因子: 56.9
作者:
Henikoff, S;Greene, EA;Hood, L
通讯作者: Hood, L
DOI: 10.1016/s0097-8485(99)00011-x
发表时间: 1999-01-01
期刊: COMPUTERS & CHEMISTRY
影响因子: --
作者:
Gouzy, J;Corpet, F;Kahn, D
通讯作者: Kahn, D