Porting and optimizing UniFrac for GPUs

Porting and optimizing UniFrac for GPUs
复制标题

为 GPU 移植和优化 UniFrac

DOI:
--
复制
发表时间:
2020
期刊:
arXiv.org
影响因子:
--
通讯作者:
R. Knight
R. Knight
中科院分区:
--
文献类型:
--
作者:
I. Sfiligoi;Daniel McDonald;R. Knight

文献摘要

参考文献

被引文献

相似文献

UniFrac是微生物组研究中常用的指标,用于相互比较微生物组概况(“β多样性”)。最近实施的条纹UniFrac添加的能力,分裂成许多独立的子问题,并表现出近线性缩放的问题。在本文中,我们描述了将Striped Unifrac移植和优化到GPU的步骤。我们将在已发布的地球微生物组项目数据集上计算UniFrac的运行时间从Intel Xeon E5-2680 v4 CPU上的13小时缩短到NVIDIA Tesla V100 GPU上的12分钟,以及NVIDIA GTX 1050笔记本电脑上的约1小时(精度略有损失)。在包含113 k个样本的较大数据集上计算UniFrac,将CPU上的运行时间从一个多月缩短到V100上的不到2小时,以及NVIDIA RTX 2080 TI GPU上的9小时(精度略有损失)。这是通过使用OpenACC生成GPU卸载代码和改进内存访问模式来实现的。一个BSD许可的实现是可用的,它产生一个由任何编程语言支持的C共享库。
UniFrac is a commonly used metric in microbiome research for comparing microbiome profiles to one another ("beta diversity"). The recently implemented Striped UniFrac added the capability to split the problem into many independent subproblems and exhibits near linear scaling. In this paper we describe steps undertaken in porting and optimizing Striped Unifrac to GPUs. We reduced the run time of computing UniFrac on the published Earth Microbiome Project dataset from 13 hours on an Intel Xeon E5-2680 v4 CPU to 12 minutes on an NVIDIA Tesla V100 GPU, and to about one hour on a laptop with NVIDIA GTX 1050 (with minor loss in precision). Computing UniFrac on a larger dataset containing 113k samples reduced the run time from over one month on the CPU to less than 2 hours on the V100 and 9 hours on an NVIDIA RTX 2080TI GPU (with minor loss in precision). This was achieved by using OpenACC for generating the GPU offload code and by improving the memory access patterns. A BSD-licensed implementation is available, which produces a C shared library linkable by any programming language.
DOI: 10.1038/nature24644
发表时间: 2017-11-23
期刊: Nature
影响因子: 64.8
作者:
Gaudelli NM;Komor AC;Rees HA;Packer MS;Badran AH;Bryson DI;Liu DR
通讯作者: Liu DR