Accelerating read mapping with FastHASH.

Accelerating read mapping with FastHASH.
复制标题

DOI:
10.1186/1471-2164-14-s1-s13
复制
发表时间:
2013
期刊:
影响因子:
4.4
通讯作者:
Alkan C
Alkan C
中科院分区:
生物学2区
文献类型:
--
作者:
Xin H;Lee D;Hormozdiari F;Yedkar S;Mutlu O;Alkan C

文献摘要

被引文献

相似文献

随着下一代测序(NGS)技术的引入,我们面临着基因组序列数据量的指数级增长。下一代测序的所有医学和遗传学应用的成功关键依赖于能够快速准确地处理和分析海量序列数据的计算技术的存在。遗憾的是,目前的读映射算法在处理NGS产生的海量数据方面存在困难。我们提出了一种新的算法FastHASH,该算法在保持基于种子和扩展型哈希表的读映射算法的高灵敏度和全面性的同时,大大提高了该算法的性能。FastHASH是一种与所有种子和扩展类读取映射算法兼容的通用算法。它介绍了两种主要技术,即邻接过滤和廉价的K-mer选择。我们实现了FastHASH并将其合并到流行的Read映射程序mrFAST的代码库中。根据编辑距离的不同,我们观察到了高达19倍的加速,同时仍然保持了100%的敏感度和高度的综合性。
With the introduction of next-generation sequencing (NGS) technologies, we are facing an exponential increase in the amount of genomic sequence data. The success of all medical and genetic applications of next-generation sequencing critically depends on the existence of computational techniques that can process and analyze the enormous amount of sequence data quickly and accurately. Unfortunately, the current read mapping algorithms have difficulties in coping with the massive amounts of data generated by NGS. We propose a new algorithm, FastHASH, which drastically improves the performance of the seed-and-extend type hash table based read mapping algorithms, while maintaining the high sensitivity and comprehensiveness of such methods. FastHASH is a generic algorithm compatible with all seed-and-extend class read mapping algorithms. It introduces two main techniques, namely Adjacency Filtering, and Cheap K-mer Selection. We implemented FastHASH and merged it into the codebase of the popular read mapping program, mrFAST. Depending on the edit distance cutoffs, we observed up to 19-fold speedup while still maintaining 100% sensitivity and high comprehensiveness.