The parallelism motifs of genomic data analysis

The parallelism motifs of genomic data analysis
复制标题

DOI:
10.1098/rsta.2019.0394
复制
发表时间:
2020-01
期刊:
Philosophical transactions. Series A, Mathematical, physical, and engineering sciences
影响因子:
--
通讯作者:
K. Yelick;A. Buluç;M. Awan;A. Azad;Benjamin Brock;R. Egan;S. Ekanayake;Marquita Ellis;E. Georganas;Giulia Guidi;S. Hofmeyr;Oguz Selvitopi;Cristina Teodoropol;L. Oliker
K. Yelick;A. Buluç;M. Awan;A. Azad;Benjamin Brock;R. Egan;S. Ekanayake;Marquita Ellis;E. Georganas;Giulia Guidi;S. Hofmeyr;Oguz Selvitopi;Cristina Teodoropol;L. Oliker
中科院分区:
其他
文献类型:
--
作者:
K. Yelick;A. Buluç;M. Awan;A. Azad;Benjamin Brock;R. Egan;S. Ekanayake;Marquita Ellis;E. Georganas;Giulia Guidi;S. Hofmeyr;Oguz Selvitopi;Cristina Teodoropol;L. Oliker

文献摘要

被引文献

相似文献

随着测序成本的持续下降和小型测序设备的可用,基因组数据集正在急剧增长。庞大的社区数据库存储并与研究社区共享这些数据,但其中一些基因组数据分析问题需要大规模计算平台来满足内存和计算需求。这些应用程序不同于当今在高端并行系统上占主导地位的工作量的科学模拟,并且对编程支持,软件库和并行架构设计提出了不同的要求。例如,它们涉及不规则的通信模式,例如对共享数据结构的异步更新。我们考虑了高性能基因组学分析中的几个问题,包括单基因组和宏基因组的比对、分析、聚类和组装。我们确定了一些常见的计算模式或“图案”,有助于通知并行化策略,并比较我们的图案的一些既定的列表,认为至少有两个关键的模式,排序和哈希,是失踪。这篇文章是讨论会议的一部分问题“高性能计算科学的数值算法”。
Genomic datasets are growing dramatically as the cost of sequencing continues to decline and small sequencing devices become available. Enormous community databases store and share these data with the research community, but some of these genomic data analysis problems require large-scale computational platforms to meet both the memory and computational requirements. These applications differ from scientific simulations that dominate the workload on high-end parallel systems today and place different requirements on programming support, software libraries and parallel architectural design. For example, they involve irregular communication patterns such as asynchronous updates to shared data structures. We consider several problems in high-performance genomics analysis, including alignment, profiling, clustering and assembly for both single genomes and metagenomes. We identify some of the common computational patterns or ‘motifs’ that help inform parallelization strategies and compare our motifs to some of the established lists, arguing that at least two key patterns, sorting and hashing, are missing. This article is part of a discussion meeting issue ‘Numerical algorithms for high-performance computational science’.