Parallelization of local BLAST service on workstation clusters

Parallelization of local BLAST service on workstation clusters
复制标题

DOI:
10.1016/s0167-739x(00)00057-1
复制
发表时间:
2001-04-01
期刊:
FUTURE GENERATION COMPUTER SYSTEMS
影响因子:
--
通讯作者:
Roberts, CA
Roberts, CA
中科院分区:
其他
文献类型:
--
作者:
Braun, RC;Pedretti, KT;Roberts, CA

文献摘要

被引文献

相似文献

本文描述了提高人类基因组计划(HGP)中最常见和日益重要的方面之一的性能的方法-大容量,批量比较DNA序列数据。这种基本的比较操作通常由著名的BLAST程序对一个主题序列与国际上可用的近500万个目标序列数据库进行比较,每天已经被世界各地的研究人员使用数十万次。目前,它仍然主要用于单个查询或小批量查询模式。随着整个人类基因组序列的接近完成,功能基因组学领域和基因微阵列的使用正在崭露头角。这些发展将需要更有效的数据处理方法,这将使单处理器在功能强大的工作站上实现变得不可行的。我们描述了BLAST的三个主要并行组件。第一个是序列到序列的比较级别。第二种方法是跨分区和分布式数据库并行处理单个查询。最后,查询集本身在一组具有复制或分区数据库的服务器上进行分区。这三种方法可以单独使用,也可以协同使用。我们当前的实现描述了并行处理批处理请求,我们对其他级别的实现计划也进行了描述。研究结果最终将应用于这种即将成为原始计算机操作的硬件辅助。(C) 2001 Elsevier Science B.V.版权所有
This paper describes approaches to improve the performance of one of the most common and increasingly important aspects of the Human Genome Project (HGP)-large-volume, batch comparison of DNA sequence data. This basic comparison operation, usually carried out by the well-known BLAST program on one subject sequence against the internationally available databases of nearly five million target sequences, is already used hundreds of thousands of times each day by researchers around the world. At present, it is still used primarily in single query, or small batch query mode. As the entire sequence of the human genome nears completion, the area of functional genomics, and the use of micro-arrays of sets of genes, is coming to the fore. These developments will demand ever more efficient means of BLASTing sets of data that will make single processor implementation on powerful workstations infeasible. We describe the three primary parallel components to BLAST. The first is at the sequence-to-sequence comparison level. The second parallelizes a single query across a partitioned and distributed database. Finally, the set of queries themselves are partitioned across a set of servers with replicated or partitioned databases. The three methods may be employed alone or in concert. Our current implementation is described which parallelizes batch requests, and our plans for implementation of the other levels is also described. The results will ultimately be applied to hardware assistance for this soon-to-be primitive computer operation. (C) 2001 Elsevier Science B.V. All rights reserved.