Microbase2.0: A Generic Framework for Computationally Intensive Bioinformatics Workflows in the Cloud

Microbase2.0: A Generic Framework for Computationally Intensive Bioinformatics Workflows in the Cloud
复制标题

DOI:
10.1515/jib-2012-212
复制
发表时间:
2012
影响因子:
1.9
通讯作者:
Keith Flanagan;S. Nakjang;J. Hallinan;C. Harwood;R. Hirt;M. Pocock;A. Wipat
Keith Flanagan;S. Nakjang;J. Hallinan;C. Harwood;R. Hirt;M. Pocock;A. Wipat
中科院分区:
--
文献类型:
--
作者:
Keith Flanagan;S. Nakjang;J. Hallinan;C. Harwood;R. Hirt;M. Pocock;A. Wipat

文献摘要

被引文献

相似文献

摘要随着生物信息学数据集变得越来越大,分析变得越来越复杂,需要数据处理基础设施来跟上技术的发展。一种解决方案是应用网格和云技术来解决分析高吞吐量数据集的计算需求。我们提出了一种方法,用于编写新的,或包装现有的应用程序,和一个参考实现的框架,Microbase2.0,使用网格和云技术执行这些应用程序。我们使用Microbase2.0开发了一个自动化的基于云的生物信息学工作流,该工作流在两个不同的Amazon EC2数据中心和纽卡斯尔大学Condor Grid上同时执行。这个系统在不到两个月的时间里完成了几个CPU年的计算工作。该工作流程产生了一个详细的数据集,描述了来自867个分类群的3,021,490种蛋白质的细胞定位,包括细菌,古细菌和单细胞真核生物。Microbase2.0可从http://www.microbase.org.uk/免费获得。
Summary As bioinformatics datasets grow ever larger, and analyses become increasingly complex, there is a need for data handling infrastructures to keep pace with developing technology. One solution is to apply Grid and Cloud technologies to address the computational requirements of analysing high throughput datasets. We present an approach for writing new, or wrapping existing applications, and a reference implementation of a framework, Microbase2.0, for executing those applications using Grid and Cloud technologies. We used Microbase2.0 to develop an automated Cloud-based bioinformatics workflow executing simultaneously on two different Amazon EC2 data centres and the Newcastle University Condor Grid. Several CPU years’ worth of computational work was performed by this system in less than two months. The workflow produced a detailed dataset characterising the cellular localisation of 3,021,490 proteins from 867 taxa, including bacteria, archaea and unicellular eukaryotes. Microbase2.0 is freely available from http://www.microbase.org.uk/.