Galaxy CloudMan: delivering cloud compute clusters.

Galaxy CloudMan: delivering cloud compute clusters.
复制标题

DOI:
10.1186/1471-2105-11-s12-s4
复制
发表时间:
2010-12-21
期刊:
影响因子:
3
通讯作者:
Taylor J
Taylor J
中科院分区:
生物学4区
文献类型:
--
作者:
Afgan E;Baker D;Coraor N;Chapman B;Nekrutenko A;Taylor J

文献摘要

被引文献

相似文献

高通量测序的广泛采用大大增加了进行基因组研究所需的计算基础设施的规模和复杂性。建设和维护本地基础设施的一个替代办法是“云计算”,原则上,云计算可按需提供灵活的计算基础设施。然而,云计算资源还不适合实验生物学家立即“原样”使用。我们提出了一个云资源管理系统,使个人研究人员可以组成和控制亚马逊的EC2云基础设施上的任意大小的计算集群,而没有任何信息学要求。在该系统中,NERC Bio-Linux团队打包的整套生物工具(http://nebc.nerc.ac.uk/tools/bio-linux)可立即使用。所提供的解决方案可以仅使用Web浏览器创建一个完全配置的计算集群,以便在不到五分钟的时间内执行分析。此外,我们还提供了一种自动化的方法来构建云资源的自定义部署。这种方法提高了结果的可重复性,并且如果需要,允许个人和实验室添加或定制其他可用的云系统,以更好地满足他们的需求。在Amazon EC2云中部署计算集群所需的知识和相关工作量并不小。本文提出的解决方案消除了这些障碍,使研究人员能够部署他们所需的计算能力,结合大量现有的分析软件,以处理持续的数据洪流。
Widespread adoption of high-throughput sequencing has greatly increased the scale and sophistication of computational infrastructure needed to perform genomic research. An alternative to building and maintaining local infrastructure is “cloud computing”, which, in principle, offers on demand access to flexible computational infrastructure. However, cloud computing resources are not yet suitable for immediate “as is” use by experimental biologists. We present a cloud resource management system that makes it possible for individual researchers to compose and control an arbitrarily sized compute cluster on Amazon’s EC2 cloud infrastructure without any informatics requirements. Within this system, an entire suite of biological tools packaged by the NERC Bio-Linux team (http://nebc.nerc.ac.uk/tools/bio-linux) is available for immediate consumption. The provided solution makes it possible, using only a web browser, to create a completely configured compute cluster ready to perform analysis in less than five minutes. Moreover, we provide an automated method for building custom deployments of cloud resources. This approach promotes reproducibility of results and, if desired, allows individuals and labs to add or customize an otherwise available cloud system to better meet their needs. The expected knowledge and associated effort with deploying a compute cluster in the Amazon EC2 cloud is not trivial. The solution presented in this paper eliminates these barriers, making it possible for researchers to deploy exactly the amount of computing power they need, combined with a wealth of existing analysis software, to handle the ongoing data deluge.