MultiPhyl: a high-throughput phylogenomics webserver using distributed computing

MultiPhyl: a high-throughput phylogenomics webserver using distributed computing
复制标题

DOI:
10.1093/nar/gkm359
复制
发表时间:
2007-07-01
影响因子:
14.9
通讯作者:
McInerney, James O.
McInerney, James O.
中科院分区:
生物学2区
文献类型:
--
作者:
Keane, Thomas M.;Naughton, Thomas J.;McInerney, James O.

文献摘要

被引文献

相似文献

随着全测序基因组数量的稳步增加,人们对从大量单个基因家族中进行大规模的系统发育分析越来越感兴趣。最大似然法(ML)已多次被证明是构建系统发育最准确的方法之一。最近,基于最大似然法的树搜索方法在算法上有了一些改进。然而,使用一台计算机分析许多基因家族的进化史仍然需要很长时间。分布式计算是指组合多台计算机的计算能力以便执行一些更大的整体计算的方法。在这篇文章中,我们提出了第一个高吞吐量的分布式系统发育平台MultiPhyl的实现,该平台能够利用许多不同的非专用机器的空闲计算资源来形成一个系统发育超级计算机。MultiPhyl允许用户同时上传数百或数千个氨基酸或核苷酸比对,并使用许多台式机执行计算密集型任务,如模型选择、树搜索和每个比对的引导。该程序实现了一套88个氨基酸模型和56个核苷酸最大似然模型,以及用于在可选模型之间进行选择的各种统计方法。公众可通过以下网址获得MultiPhyl网络服务器:http://www.cs.nuim.ie/distributed/multiphyl.php.
With the number of fully sequenced genomes increasing steadily, there is greater interest in performing large-scale phylogenomic analyses from large numbers of individual gene families. Maximum likelihood (ML) has been shown repeatedly to be one of the most accurate methods for phylogenetic construction. Recently, there have been a number of algorithmic improvements in maximum-likelihood-based tree search methods. However, it can still take a long time to analyse the evolutionary history of many gene families using a single computer. Distributed computing refers to a method of combining the computing power of multiple computers in order to perform some larger overall calculation. In this article, we present the first high-throughput implementation of a distributed phylogenetics platform, MultiPhyl, capable of using the idle computational resources of many heterogeneous non-dedicated machines to form a phylogenetics supercomputer. MultiPhyl allows a user to upload hundreds or thousands of amino acid or nucleotide alignments simultaneously and perform computationally intensive tasks such as model selection, tree searching and boot-strapping of each of the alignments using many desktop machines. The program implements a set of 88 amino acid models and 56 nucleotide maximum likelihood models and a variety of statistical methods for choosing between alternative models. A MultiPhyl webserver is available for public use at: http://www.cs.nuim.ie/distributed/multiphyl.php.