NGSANE: a lightweight production informatics framework for high-throughput data analysis

NGSANE: a lightweight production informatics framework for high-throughput data analysis
复制标题

DOI:
10.1093/bioinformatics/btu036
复制
发表时间:
2014-05-15
期刊:
影响因子:
5.8
通讯作者:
Bauer, Denis C.
Bauer, Denis C.
中科院分区:
生物学3区
文献类型:
--
作者:
Buske, Fabian A.;French, Hugh J.;Bauer, Denis C.

文献摘要

被引文献

相似文献

分析下一代测序数据的初始步骤可以通过软件“流水线”的方式实现自动化。然而,由于不断发展的技术和分析方法,个别组件迅速贬值,往往导致生产信息管道的整个版本过时。从Linux bash命令构建管道可以使用热插拔模块化组件,而不是使用更严格的高级语言程序调用包装,这是在类似的已发布的流水线系统中实现的。在这里,我们提出了下一代企业排序分析(NGSANE),这是一个基于Linux的、支持高性能计算的框架,它最大限度地减少了新项目的设置和处理开销,同时在处理原始序列数据时保持了定制脚本的完全灵活性。
The initial steps in the analysis of next-generation sequencing data can be automated by way of software 'pipelines'. However, individual components depreciate rapidly because of the evolving technology and analysis methods, often rendering entire versions of production informatics pipelines obsolete. Constructing pipelines from Linux bash commands enables the use of hot swappable modular components as opposed to the more rigid program call wrapping by higher level languages, as implemented in comparable published pipelining systems.Here we present Next Generation Sequencing ANalysis for Enterprises (NGSANE), a Linux-based, high-performance-computing-enabled framework that minimizes overhead for set up and processing of new projects, yet maintains full flexibility of custom scripting when processing raw sequence data.