Rapid Likelihood Analysis on Large Phylogenies Using Partial Sampling of Substitution Histories

Rapid Likelihood Analysis on Large Phylogenies Using Partial Sampling of Substitution Histories
复制标题

DOI:
10.1093/molbev/msp228
复制
发表时间:
2010-02-01
影响因子:
10.7
通讯作者:
Pollock, David D.
Pollock, David D.
中科院分区:
生物学1区
文献类型:
--
作者:
de Koning, A. P. Jason;Gu, Wanjun;Pollock, David D.

文献摘要

被引文献

相似文献

基于似然的方法可以从更大的数据集更详细、更精确地重建进化过程。目前正在产生的极其庞大的比较基因组数据集为理解分子进化创造了新的机会,但分析如此大量的数据带来了不断升级的计算挑战。最近开发的马尔可夫链蒙特卡罗方法增加替代历史是一个有希望的方法来减轻这些计算成本。我们分析了几种此类方法的计算成本,并考虑了它们如何随模型和数据集复杂性进行扩展。这为理解最重要的计算瓶颈提供了一个理论框架,引导我们将条件路径集成方法的新变体与其他人的最新进展结合起来。由此产生的技术(替代历史的“部分抽样”)比我们考虑的所有其他方法都要快得多。它是准确的,易于实现的,并且在模型复杂性和数据集大小的维度上非常好地扩展。特别是,使用新方法对未观察到的替代历史进行采样的时间复杂度比现有方法快得多,并且模型参数和分支长度更新与数据集大小无关。我们比较了224种方法在哺乳动物细胞色素-b序列上的表现。对于一个简单的核苷酸替换模型,部分采样比连续时间内采样的PhyloBayes程序快至少10倍,比使用完全集成的替换历史时快100倍。在一般可逆的氨基酸取代模型下,部分采样方法比使用完全集成的取代历史时快1600倍,证实了模型状态空间复杂性的显着改善。因此,替换的部分抽样极大地提高了似然方法在分析大型数据集上复杂进化过程的效用。
Likelihood-based approaches can reconstruct evolutionary processes in greater detail and with better precision from larger data sets. The extremely large comparative genomic data sets that are now being generated thus create new opportunities for understanding molecular evolution, but analysis of such large quantities of data poses escalating computational challenges. Recently developed Markov chain Monte Carlo methods that augment substitution histories are a promising approach to alleviate these computational costs. We analyzed the computational costs of several such approaches, considering how they scale with model and data set complexity. This provided a theoretical framework to understand the most important computational bottlenecks, leading us to combine novel variations of our conditional pathway integration approach with recent advances made by others. The resulting technique ("partial sampling" of substitution histories) is considerably faster than all other approaches we considered. It is accurate, simple to implement, and scales exceptionally well with dimensions of model complexity and data set size. In particular, the time complexity of sampling unobserved substitution histories using the new method is much faster than previously existing methods, and model parameter and branch length updates are independent of data set size. We compared the performance of methods on a 224-taxon set of mammalian cytochrome-b sequences. For a simple nucleotide substitution model, partial sampling was at least 10 times faster than the PhyloBayes program, which samples substitutions in continuous time, and about 100 times faster than when using fully integrated substitution histories. Under a general reversible model of amino acid substitution, the partial sampling method was 1,600 times faster than when using fully integrated substitution histories, confirming significantly improved scaling with model state-space complexity. Partial sampling of substitutions thus dramatically improves the utility of likelihood approaches for analyzing complex evolutionary processes on large data sets.