Optimizing High Performance Distributed Memory Parallel Hash Tables for DNA k-mer Counting
Optimizing High Performance Distributed Memory Parallel Hash Tables for DNA k-mer Counting
复制标题
优化用于 DNA k 聚体计数的高性能分布式内存并行哈希表
DOI:
10.1109/sc.2018.00014
复制
发表时间:
2018
期刊:
影响因子:
--
通讯作者:
S. Aluru
中科院分区:
文献类型:
--
作者:
Tony Pan;Sanchit Misra;S. Aluru
High-throughput DNA sequencing is the mainstay of modern genomics research. A common operation used in bioinformatic analysis for many applications of high-throughput sequencing is the counting and indexing of fixed length substrings of DNA sequences called k-mers. Counting k-mers is often accomplished via hashing, and distributed memory k-mer counting algorithms for large datasets are memory access and network communication bound. In this work, we present two optimized distributed parallel hash table techniques that utilize cache friendly algorithms for local hashing, overlapped communication and computation to hide communication costs, and vectorized hash functions that are specialized for fc-mer and other short key indices. On 4096 cores of the NERSC Cori supercomputer, our implementation completed index construction and query on an approximately 1 TB human genome dataset in just 11.8 seconds and 5.8 seconds, demonstrating speedups of 2.06× and 3.7×, respectively, over the previous state-of-the-art distributed memory k-mer counter.
影响因子:
12
作者:
Koepfli KP;Paten B;Genome 10K Community of Scientists;O'Brien SJ
通讯作者:
O'Brien SJ
影响因子:
7
作者:
Salzberg, Steven L.;Phillippy, Adam M.;Yorke, James A.
通讯作者:
Yorke, James A.