GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks

GraphPIM: Enabling Instruction-Level PIM Offloading in Graph Computing Frameworks
复制标题

DOI:
10.1109/hpca.2017.54
复制
发表时间:
2017-02
期刊:
2017 IEEE International Symposium on High Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Lifeng Nai;Ramyad Hadidi;Jaewoong Sim;Hyojong Kim;Pranith Kumar;Hyesoon Kim
Lifeng Nai;Ramyad Hadidi;Jaewoong Sim;Hyojong Kim;Pranith Kumar;Hyesoon Kim
中科院分区:
其他
文献类型:
--
作者:
Lifeng Nai;Ramyad Hadidi;Jaewoong Sim;Hyojong Kim;Pranith Kumar;Hyesoon Kim

文献摘要

被引文献

相似文献

随着数据科学的出现,图计算现在变得越来越重要。不幸的是,由于执行原子操作的开销和存储器子系统的低效利用,图计算在映射到现代计算系统时通常遭受较差的性能。同时,诸如混合存储器立方体(HMC)之类的新兴技术使得能够在指令级利用卸载操作来实现存储器中处理(PIM)功能。将指令卸载到PIM侧对于克服图计算的性能瓶颈具有相当大的潜力。尽管如此,图形工作负载的此功能尚未得到充分探索,其应用程序和缺点迄今尚未得到很好的识别。在本文中,我们提出了GraphPIM,一个全栈的解决方案,图计算,实现更高的性能,使用PIM功能。我们对现代图形工作负载进行分析,以评估PIM卸载的适用性,并提出硬件和软件机制,以有效地利用PIM功能。遵循真实世界的HMC 2.0规范,GraphPIM为图形应用程序提供了性能优势,而无需任何用户代码修改或伊萨更改。此外,我们提出了一个扩展PIM操作,可以进一步带来更多的图形应用程序的性能优势。评估结果表明,GraphPIM实现了高达2.4倍的加速比,能耗降低了37%。
With the emergence of data science, graph computing has become increasingly important these days. Unfortunately, graph computing typically suffers from poor performance when mapped to modern computing systems because of the overhead of executing atomic operations and inefficient utilization of the memory subsystem. Meanwhile, emerging technologies, such as Hybrid Memory Cube (HMC), enable the processing-in-memory (PIM) functionality with offloading operations at an instruction level. Instruction offloading to the PIM side has considerable potentials to overcome the performance bottleneck of graph computing. Nevertheless, this functionality for graph workloads has not been fully explored, and its applications and shortcomings have not been well identified thus far. In this paper, we present GraphPIM, a full-stack solution for graph computing that achieves higher performance using PIM functionality. We perform an analysis on modern graph workloads to assess the applicability of PIM offloading and present hardware and software mechanisms to efficiently make use of the PIM functionality. Following the real-world HMC 2.0 specification, GraphPIM provides performance benefits for graph applications without any user code modification or ISA changes. In addition, we propose an extension to PIM operations that can further bring performance benefits for more graph applications. The evaluation results show that GraphPIM achieves up to a 2.4X speedup with a 37% reduction in energy consumption.