SeqWare Query Engine: storing and searching sequence data in the cloud.

SeqWare Query Engine: storing and searching sequence data in the cloud.
复制标题

DOI:
10.1186/1471-2105-11-s12-s2
复制
发表时间:
2010-12-21
期刊:
影响因子:
3
通讯作者:
Nelson SF
Nelson SF
中科院分区:
生物学4区
文献类型:
--
作者:
O'Connor BD;Merriman B;Nelson SF

文献摘要

被引文献

相似文献

自从下一代DNA测序仪问世以来,测序仪通量的快速增加以及相关的成本下降,导致在过去几年中对十几个人类基因组进行了重新测序。这些努力仅仅是一个前奏曲,在未来,基因组重测序将是司空见惯的生物医学研究和临床应用。测序仪输出的急剧增加使计算基础设施的各个方面都变得紧张,特别是数据库和查询接口。云计算的出现,以及各种旨在处理千万亿次数据集的强大工具,为这些不断增长的需求提供了令人信服的解决方案。在这项工作中,我们提出了SeqWare查询引擎,它是使用现代云计算技术创建的,旨在支持来自数千个基因组的数据库信息。我们的后端实现是使用Hadoop项目中高度可扩展的NoSQL HBase数据库构建的。我们还创建了一个基于Web的前端,提供了一个编程和交互式查询界面,并与广泛使用的基因组浏览器和工具集成。使用查询引擎,用户可以加载和查询变体(SNV,indel,translocations等),并提供丰富的注释,包括覆盖范围和功能后果。作为概念证明,我们加载了包括U87 MG细胞系在内的几个全基因组数据集。我们还使用了一个多形性胶质母细胞瘤肿瘤/正常对来分析性能,并提供了在查询引擎中使用Hadoop MapReduce框架的示例。该软件是开放源代码的,可从SeqWare项目(http://www.example.com)免费获得。seqware.sourceforge.net SeqWare查询引擎提供了一种简单的方法,使程序员和非程序员都可以访问U87 MG基因组。这使得能够更快和更开放地探索结果,更快地调整启发式变异识别过滤器的参数,以及简化分析工具开发的通用数据接口。支持的数据类型范围,查询和与现有工具集成的便利性,以及基于云的底层技术的强大可扩展性,使SeqWare查询引擎非常适合存储和搜索不断增长的基因组序列数据集。
Since the introduction of next-generation DNA sequencers the rapid increase in sequencer throughput, and associated drop in costs, has resulted in more than a dozen human genomes being resequenced over the last few years. These efforts are merely a prelude for a future in which genome resequencing will be commonplace for both biomedical research and clinical applications. The dramatic increase in sequencer output strains all facets of computational infrastructure, especially databases and query interfaces. The advent of cloud computing, and a variety of powerful tools designed to process petascale datasets, provide a compelling solution to these ever increasing demands. In this work, we present the SeqWare Query Engine which has been created using modern cloud computing technologies and designed to support databasing information from thousands of genomes. Our backend implementation was built using the highly scalable, NoSQL HBase database from the Hadoop project. We also created a web-based frontend that provides both a programmatic and interactive query interface and integrates with widely used genome browsers and tools. Using the query engine, users can load and query variants (SNVs, indels, translocations, etc) with a rich level of annotations including coverage and functional consequences. As a proof of concept we loaded several whole genome datasets including the U87MG cell line. We also used a glioblastoma multiforme tumor/normal pair to both profile performance and provide an example of using the Hadoop MapReduce framework within the query engine. This software is open source and freely available from the SeqWare project (http://seqware.sourceforge.net). The SeqWare Query Engine provided an easy way to make the U87MG genome accessible to programmers and non-programmers alike. This enabled a faster and more open exploration of results, quicker tuning of parameters for heuristic variant calling filters, and a common data interface to simplify development of analytical tools. The range of data types supported, the ease of querying and integrating with existing tools, and the robust scalability of the underlying cloud-based technologies make SeqWare Query Engine a nature fit for storing and searching ever-growing genome sequence datasets.