SS-Wrapper: a package of wrapper applications for similarity searches on Linux clusters.

SS-Wrapper: a package of wrapper applications for similarity searches on Linux clusters.
复制标题

DOI:
10.1186/1471-2105-5-171
复制
发表时间:
2004-10-28
期刊:
影响因子:
3
通讯作者:
Lefkowitz EJ
Lefkowitz EJ
中科院分区:
生物学4区
文献类型:
--
作者:
Wang C;Lefkowitz EJ

文献摘要

相似文献

大规模序列比较是现代分子生物学中进行生物推断的有力工具。将新序列与注释数据库中的序列进行比较是关于这些序列的功能和结构信息的有用来源。使用软件,如基本的本地比对搜索工具(BLAST)或HMMPFAM,以确定新测序的遗传物质片段和数据库中的那些之间的统计学显着匹配是大多数分子生物学家的一项重要任务。搜索算法本质上是缓慢的和数据密集型的,特别是鉴于由于高通量DNA测序技术的出现而导致的生物序列数据库的快速增长。因此,传统的生物信息学工具在PC上甚至在专用UNIX服务器上都是不切实际的。为了利用更大的数据库和更可靠的方法,高性能计算变得必要。我们描述了SS-包装器(相似性搜索包装器),包装器应用程序,可以并行相似性搜索应用程序在Linux集群的实现。我们的包装器利用查询分割搜索(QS搜索)方法来并行化序列数据库搜索应用程序。它考虑了群集上每个节点之间的负载平衡,以最大限度地利用资源。QS-search旨在使用相同的界面包装许多不同的搜索工具,例如BLAST和HMMPFAM。这种实现不会改变原始程序,因此新获得的程序和程序更新应该很容易适应。使用QS搜索优化BLAST和HMMPFAM的基准实验表明,QS搜索几乎线性地加速了这些程序的性能与所使用的CPU数量成比例。我们还实现了一个包装器,利用数据库分割方法(DS-BLAST),提供了一个互补的解决方案,BLAST搜索时,数据库太大,适合到一个单一的节点的内存。QS-search和DS-BLAST结合使用,提供了一种灵活的解决方案,以适应高性能计算环境中的顺序相似性搜索应用。它们的易用性和包装各种数据库搜索程序的能力提供了一个分析架构,以帮助经验丰富的生物信息学家和湿板凳生物学家。
Large-scale sequence comparison is a powerful tool for biological inference in modern molecular biology. Comparing new sequences to those in annotated databases is a useful source of functional and structural information about these sequences. Using software such as the basic local alignment search tool (BLAST) or HMMPFAM to identify statistically significant matches between newly sequenced segments of genetic material and those in databases is an important task for most molecular biologists. Searching algorithms are intrinsically slow and data-intensive, especially in light of the rapid growth of biological sequence databases due to the emergence of high throughput DNA sequencing techniques. Thus, traditional bioinformatics tools are impractical on PCs and even on dedicated UNIX servers. To take advantage of larger databases and more reliable methods, high performance computation becomes necessary. We describe the implementation of SS-Wrapper (Similarity Search Wrapper), a package of wrapper applications that can parallelize similarity search applications on a Linux cluster. Our wrapper utilizes a query segmentation-search (QS-search) approach to parallelize sequence database search applications. It takes into consideration load balancing between each node on the cluster to maximize resource usage. QS-search is designed to wrap many different search tools, such as BLAST and HMMPFAM using the same interface. This implementation does not alter the original program, so newly obtained programs and program updates should be accommodated easily. Benchmark experiments using QS-search to optimize BLAST and HMMPFAM showed that QS-search accelerated the performance of these programs almost linearly in proportion to the number of CPUs used. We have also implemented a wrapper that utilizes a database segmentation approach (DS-BLAST) that provides a complementary solution for BLAST searches when the database is too large to fit into the memory of a single node. Used together, QS-search and DS-BLAST provide a flexible solution to adapt sequential similarity searching applications in high performance computing environments. Their ease of use and their ability to wrap a variety of database search programs provide an analytical architecture to assist both the seasoned bioinformaticist and the wet-bench biologist.