Agile parallel bioinformatics workflow management using Pwrake.

Agile parallel bioinformatics workflow management using Pwrake.
复制标题

使用Pwrake的敏捷并行生物信息学工作流程管理。

DOI:
10.1186/1756-0500-4-331
复制
发表时间:
2011-09-08
期刊:
影响因子:
1.8
通讯作者:
Yoshiura K
Yoshiura K
中科院分区:
其他
文献类型:
--
作者:
Mishima H;Sasaki K;Tanaka M;Tatebe O;Yoshiura K

文献摘要

被引文献

相似文献

在生物信息学项目中,科学工作流系统被广泛用于管理计算过程。为了满足工作流管理的需求,人们提出了功能完备的工作流系统。然而,这样的系统往往是过重的实际生物信息学的做法。我们意识到,在科学的工作流管理中,快速部署实施先进算法和数据格式的尖端软件,以及不断适应计算资源和环境的变化,往往是优先考虑的问题。通过反复试验和错误之后的迭代开发阶段,这些特性与敏捷软件开发方法具有更大的亲和力。在这里,我们展示了科学工作流系统Pwrake在生物信息学工作流中的应用。Pwrake是Ruby标准构建工具Rake的并行工作流扩展,其灵活性已在天文学领域得到证明。因此,我们假设Pwrake在实际的生物信息学工作流程中也具有优势。我们实施了Pwrake工作流程,使用基因组分析工具包(GATK)和Dindel处理下一代测序数据。GATK和Dindel工作流分别是顺序工作流和并行工作流的典型示例。我们发现,在实践中,实际的科学工作流开发迭代两个阶段,工作流定义阶段和参数调整阶段。我们引入了单独的工作流定义来帮助关注两个开发阶段中的每一个,以及帮助器方法来简化描述。这种方法提高了迭代开发的效率。此外,我们实现了组合工作流程,以展示GATK和Dindel工作流程的模块化。Pwrake使生物信息学领域的科学工作流程的敏捷管理成为可能。基于Ruby的内部领域特定语言设计为编写科学工作流提供了rakefiles的灵活性。此外,rakefiles的可读性和可维护性可以促进科学界之间共享工作流程。GATK和Dindel的工作流程可在http://github.com/misshie/Workflows上找到。
In bioinformatics projects, scientific workflow systems are widely used to manage computational procedures. Full-featured workflow systems have been proposed to fulfil the demand for workflow management. However, such systems tend to be over-weighted for actual bioinformatics practices. We realize that quick deployment of cutting-edge software implementing advanced algorithms and data formats, and continuous adaptation to changes in computational resources and the environment are often prioritized in scientific workflow management. These features have a greater affinity with the agile software development method through iterative development phases after trial and error. Here, we show the application of a scientific workflow system Pwrake to bioinformatics workflows. Pwrake is a parallel workflow extension of Ruby's standard build tool Rake, the flexibility of which has been demonstrated in the astronomy domain. Therefore, we hypothesize that Pwrake also has advantages in actual bioinformatics workflows. We implemented the Pwrake workflows to process next generation sequencing data using the Genomic Analysis Toolkit (GATK) and Dindel. GATK and Dindel workflows are typical examples of sequential and parallel workflows, respectively. We found that in practice, actual scientific workflow development iterates over two phases, the workflow definition phase and the parameter adjustment phase. We introduced separate workflow definitions to help focus on each of the two developmental phases, as well as helper methods to simplify the descriptions. This approach increased iterative development efficiency. Moreover, we implemented combined workflows to demonstrate modularity of the GATK and Dindel workflows. Pwrake enables agile management of scientific workflows in the bioinformatics domain. The internal domain specific language design built on Ruby gives the flexibility of rakefiles for writing scientific workflows. Furthermore, readability and maintainability of rakefiles may facilitate sharing workflows among the scientific community. Workflows for GATK and Dindel are available at http://github.com/misshie/Workflows.