CoPIM: A Concurrency-aware PIM Workload Offloading Architecture for Graph Applications

CoPIM: A Concurrency-aware PIM Workload Offloading Architecture for Graph Applications
复制标题

DOI:
10.1109/islped52811.2021.9502483
复制
发表时间:
2021-07
期刊:
2021 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED)
影响因子:
--
通讯作者:
Liang Yan;Mingzhe Zhang;Rujia Wang;Xiaoming Chen;Xingqi Zou;Xiaoyang Lu;Yinhe Han;Xian-He Sun
Liang Yan;Mingzhe Zhang;Rujia Wang;Xiaoming Chen;Xingqi Zou;Xiaoyang Lu;Yinhe Han;Xian-He Sun
中科院分区:
其他
文献类型:
--
作者:
Liang Yan;Mingzhe Zhang;Rujia Wang;Xiaoming Chen;Xingqi Zou;Xiaoyang Lu;Yinhe Han;Xian-He Sun

文献摘要

相似文献

内存内处理(PIM)被视为一种很有前景的解决方案,它通过尽量减少主机与内存之间的数据移动来提升图计算应用程序的性能。将哪些工作负载卸载以及如何将其卸载到PIM逻辑,这决定了PIM架构能否得到充分利用。从主机处理器向PIM端卸载过多或过少的工作负载都可能损害整体性能。另一方面,卸载粒度既要具有代表性又不能失去通用性。在本文中,我们提出了CoPIM,这是一种新颖的PIM工作负载卸载架构,它可以动态确定图工作负载的哪一部分从PIM端计算中获益更多。CoPIM聚焦于图应用程序的循环代码块,并基于并发内存访问模型评估卸载的必要性。我们还提供了详细的架构设计以支持卸载。通过这种方式,CoPIM减少了卸载指令的规模,同时以更低的能耗提升了整体性能。实验结果表明,与其他最先进的PIM工作负载卸载框架相比,CoPIM相较于PEI和GraphPIM,分别实现了19.5%和11.4%的几何平均加速比。另一方面,CoPIM相较于PEI和GraphPIM,分别平均降低了6.8%和6.5%的非核心能耗。
Processing-in-Memory (PIM) is considered a promising solution to improve the performance of graph-computing applications by minimizing the data movement between the host and memory. Which workload to offload and how to offload it to PIM logic determine whether the PIM architecture is well utilized. Offloading too much or too little workload from the host processor to the PIM side could hurt overall performance. On the other hand, the offloading granularity needs to be representative without losing generality. In this paper, we present CoPIM, a novel PIM workload offloading architecture that can dynamically determine which portion of the graph workload can benefit more from PIM-side computation. CoPIM focuses on the loop code blocks of graph applications and evaluates the necessity of offloading based on a concurrent memory access model. We also provide detailed architectural designs to support the offloading. In this way, CoPIM reduces the size of offloading instructions and also improves the overall performance with less energy consumption. The experimental results show that compared with other state-of-the-art PIM workload offloading frameworks, CoPIM achieves a speedup by the geometric mean of 19.5% and 11.4% than PEI and GraphPIM, respectively. On the other hand, CoPIM also reduces the un-core energy consumption by 6.8% and 6.5% on average over PEI and GraphPIM, respectively.