Automatic data allocation and buffer management for multi-GPU machines

Automatic data allocation and buffer management for multi-GPU machines
复制标题

多 GPU 机器的自动数据分配和缓冲区管理

DOI:
--
复制
发表时间:
2013
期刊:
ACM Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
通讯作者:
Uday Bondhugula
Uday Bondhugula
中科院分区:
--
文献类型:
--
作者:
Thejas Ramashekar;Uday Bondhugula

文献摘要

被引文献

相似文献

多GPU机器正在越来越多地用于高性能计算。这样的机器中的每个GPU都有自己的内存,并且不会与主机CPU或其他GPU共享地址空间。因此,利用多个GPU的应用程序必须在每个GPU上手动分配和管理数据。提议自动化GPU数据分配的现有作品在分配大小,利用重复使用,转移成本和可扩展性方面具有限制和效率低下。我们为多GPU机器上的仿射循环巢提供了可扩展且全自动的数据分配和缓冲管理方案。我们称其为基于边界框的内存管理器(BBMM)。 BBMM可以在运行时,在标准集合操作(​​例如联合,交叉路口)和差异,在数组数据的高偏角区域(边界框)上找到子集和超集关系。它使用这些操作以及一些编译器帮助来识别,分配和管理应用程序所需的数据,以脱节边界框。这允许它(1)完全或几乎分配到每个GPU上运行的计算所需的数据,(2)有效跟踪缓冲区分配,因此最大程度地将数据重复使用跨图块,并最大程度地将数据传输台面开销,以及(3),AS AS AS AS AS AS(3)结果,最大程度地利用了多GPU机器上组合内存的利用。 BBMM可以选择并行的转换,计算放置和调度方案,无论是静态还是动态。与当前的分配方案相比,在具有各种科学程序的四个GPU机器上进行的实验表明,BBMM将数据分配降低了75%,可产生至少88%的手动书面代码的性能,并且允许出色的弱缩放标准。
Multi-GPU machines are being increasingly used in high-performance computing. Each GPU in such a machine has its own memory and does not share the address space either with the host CPU or other GPUs. Hence, applications utilizing multiple GPUs have to manually allocate and manage data on each GPU. Existing works that propose to automate data allocations for GPUs have limitations and inefficiencies in terms of allocation sizes, exploiting reuse, transfer costs, and scalability. We propose a scalable and fully automatic data allocation and buffer management scheme for affine loop nests on multi-GPU machines. We call it the Bounding-Box-based Memory Manager (BBMM). BBMM can perform at runtime, during standard set operations like union, intersection, and difference, finding subset and superset relations on hyperrectangular regions of array data (bounding boxes). It uses these operations along with some compiler assistance to identify, allocate, and manage data required by applications in terms of disjoint bounding boxes. This allows it to (1) allocate exactly or nearly as much data as is required by computations running on each GPU, (2) efficiently track buffer allocations and hence maximize data reuse across tiles and minimize data transfer overhead, and (3) and as a result, maximize utilization of the combined memory on multi-GPU machines. BBMM can work with any choice of parallelizing transformations, computation placement, and scheduling schemes, whether static or dynamic. Experiments run on a four-GPU machine with various scientific programs showed that BBMM reduces data allocations on each GPU by up to 75% compared to current allocation schemes, yields performance of at least 88% of manually written code, and allows excellent weak scaling.