Out-of-core GPU memory management for MapReduce-based large-scale graph processing

Out-of-core GPU memory management for MapReduce-based large-scale graph processing
复制标题

DOI:
10.1109/cluster.2014.6968748
复制
发表时间:
2014-12
期刊:
2014 IEEE International Conference on Cluster Computing (CLUSTER)
影响因子:
--
通讯作者:
Koichi Shirahata;Hitoshi Sato;S. Matsuoka
Koichi Shirahata;Hitoshi Sato;S. Matsuoka
中科院分区:
其他
文献类型:
--
作者:
Koichi Shirahata;Hitoshi Sato;S. Matsuoka

文献摘要

相似文献

GPU可以加速图形处理应用程序的边缘扫描性能;然而,GPU上设备内存的容量限制了要处理的图形的大小,而处理GPU内存溢出的有效技术,包括大规模系统中的溢出检测和性能分析,还没有得到很好的研究。针对这一问题,我们提出了一种基于MapReduce的GPU核外内存管理技术,用于在基于GPU的异构型超级计算机上处理大规模图形应用。我们提出的技术通过将图形数据动态划分为多个块来自动处理来自GPU的内存溢出,并尽可能地重叠在GPU上的CPU-GPU数据传输和计算。我们在TSUBAME2.5上的1024个节点(12288个CPU核,3072个GPU)上的实验结果表明,当图形数据大小不适合在GPU上运行时,基于GPU的实现比在CPU上运行的速度快2.10倍。我们还研究了我们提出的核外GPU内存管理技术的性能特征,包括向上扩展和向外扩展方法的应用程序性能和能效。
GPUs can accelerate edge scan performance of graph processing applications; however, the capacity of device memory on GPUs limits the size of graph to process, whereas efficient techniques to handle GPU memory overflows, including overflow detection and performance analysis in large-scale systems, are not well investigated. To address the problem, we propose a MapReduce-based out-of-core GPU memory management technique for processing large-scale graph applications on heterogeneous GPU-based supercomputers. Our proposed technique automatically handles memory overflows from GPUs by dynamically dividing graph data into multiple chunks and overlaps CPU-GPU data transfer and computation on GPUs as much as possible. Our experimental results on TSUBAME2.5 using 1024 nodes (12288 CPU cores, 3072 GPUs) exhibit that our GPU-based implementation performs 2.10x faster than running on CPU when graph data size does not fit on GPUs. We also study the performance characteristics of our proposed out-of-core GPU memory management technique, including application's performance and power efficiency of scale-up and scale-out approaches.