WarpDrive: Massively Parallel Hashing on Multi-GPU Nodes

WarpDrive: Massively Parallel Hashing on Multi-GPU Nodes
复制标题

WarpDrive:多 GPU 节点上的大规模并行哈希

DOI:
--
复制
发表时间:
2018
期刊:
IEEE International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
B. Schmidt
B. Schmidt
中科院分区:
--
文献类型:
--
作者:
Daniel Jünger;Christian Hundt;B. Schmidt

文献摘要

被引文献

相似文献

哈希图是计算机科学中最通用的数据结构之一,因为其紧凑的数据布局和预期的插入和查询的恒定时间复杂性。然而,在探测阶段,关联的存储器访问模式是高度不规则的,从而导致强烈的存储器限制实现。大规模并行加速器(如支持CUDA的GPU)可以克服这一限制,因为它们的快速视频内存具有几乎1 TB/S带宽,与低于100 GB/S的最先进CPU的主内存模块相比。不幸的是,现有的单GPU哈希实现支持的哈希图的大小受到可用视频RAM数量有限的限制。因此,迫切需要构建和查询跨多个GPU的哈希图,以支持更大数据集的高速结构化存储。在本文中,我们介绍了WarpDrive-一个可扩展的、分布式的单节点多GPU实现,用于构建和查询数十亿个键-值对。我们提出了一种新的基于Subwarp的探测方案,其特征是在连续的存储区域上合并存储访问,以缓解不规则访问模式的高延迟。我们的实现在单GPU模式下实现了每秒14亿次插入,负载率为0.95,因此在P100上的性能比CUDPP库的GPU-Cuckoo实现高2.8倍。此外,我们将透明扩展到同一节点内的多个GPU,每秒最多可执行43亿次操作,从而在通过NVLink技术连接的四个P100 GPU上实现高负载率。WarpDrive是一款免费软件,可从https://github.com/sleeepyjack/warpdrive.下载
Hash maps are among the most versatile data structures in computer science because of their compact data layout and expected constant time complexity for insertion and querying. However, associated memory access patterns during the probing phase are highly irregular resulting in strongly memory-bound implementations. Massively parallel accelerators such as CUDA-enabled GPUs may overcome this limitation by virtue of their fast video memory featuring almost one TB/s bandwidth in comparison to main memory modules of state-of-the-art CPUs with less than 100 GB/s. Unfortunately, the size of hash maps supported by existing single-GPU hashing implementations is restricted by the limited amount of available video RAM. Hence, hash map construction and querying that scales across multiple GPUs is urgently needed in order to support structured storage of bigger datasets at high speeds. In this paper, we introduce WarpDrive – a scalable, distributed single-node multi-GPU implementation for the construction and querying of billions of key-value pairs. We propose a novel subwarp-based probing scheme featuring coalesced memory access over consecutive memory regions in order to mitigate the high latency of irregular access patterns. Our implementation achieves 1.4 billion insertions per second in single-GPU mode for a load factor of 0.95 thereby outperforming the GPU-cuckoo implementation of the CUDPP library by a factor of 2.8 on a P100. Furthermore, we present transparent scaling to multiple GPUs within the same node with up to 4.3 billion operations per second for high load factors on four P100 GPUs connected by NVLink technology. WarpDrive is free software and can be downloaded at https://github.com/sleeepyjack/warpdrive.