DPRml: distributed phylogeny reconstruction by maximum likelihood

DPRml: distributed phylogeny reconstruction by maximum likelihood
复制标题

DOI:
10.1093/bioinformatics/bti100
复制
发表时间:
2005-04-01
期刊:
影响因子:
5.8
通讯作者:
McCormack, GP
McCormack, GP
中科院分区:
生物学3区
文献类型:
--
作者:
Keane, TM;Naughton, TJ;McCormack, GP

文献摘要

被引文献

相似文献

动机:近年来,人们越来越有兴趣使用统计方法制作大型和准确的系统发育树。然而,对于大量的分类群,它是不可行的,以构建大而准确的树只使用一个处理器。一些专门的并行程序已经产生,试图解决巨大的计算需求的最大似然。我们表达了一些关注的并行系统发育程序,目前严重限制了广泛的可用性和使用并行计算在最大似然为基础的phylogenetic analysis.Results:我们已经确定了系统发育分析大规模异构分布式计算的适用性。我们已经完成了一个分布式和完全跨平台的系统发育树构建程序,称为分布式系统发育重建最大似然法。它使用了一个已经证明的最大似然为基础的树构建算法和流行的系统发育分析库的所有可能性计算。它提供了目前可用的最广泛的DNA替代模型之一。据我们所知,我们是第一个报告完成分布式系统发育树构建程序,可以实现近线性加速,同时只使用机器的空闲时钟周期。对于那些在学术或企业环境中拥有数百台空闲桌面计算机的人,我们已经展示了分布式计算如何提供“免费”的ML超级计算机。
Motivation: In recent years there has been increased interest in producing large and accurate phylogenetic trees using statistical approaches. However for a large number of taxa, it is not feasible to construct large and accurate trees using only a single processor. A number of specialized parallel programs have been produced in an attempt to address the huge computational requirements of maximum likelihood. We express a number of concerns about the current set of parallel phylogenetic programs which are currently severely limiting the widespread availability and use of parallel computing in maximum likelihood-based phylogenetic analysis.Results: We have identified the suitability of phylogenetic analysis to large-scale heterogeneous distributed computing. We have completed a distributed and fully cross-platform phylogenetic tree building program called distributed phylogeny reconstruction by maximum likelihood. It uses an already proven maximum likelihood-based tree building algorithm and a popular phylogenetic analysis library for all its likelihood calculations. It offers one of the most extensive sets of DNA substitution models currently available. We are the first, to our knowledge, to report the completion of a distributed phylogenetic tree building program that can achieve near-linear speedup while only using the idle clock cycles of machines. For those in an academic or corporate environment with hundreds of idle desktop machines, we have shown how distributed computing can deliver a 'free' ML supercomputer.