Accelerating Graph Computations on 3D NoC-enabled PIM Architectures

Accelerating Graph Computations on 3D NoC-enabled PIM Architectures
复制标题

加速支持 3D NoC 的 PIM 架构上的图形计算

DOI:
10.1145/3564290
复制
发表时间:
2022
影响因子:
1.4
通讯作者:
Pande, Partha Pratim
Pande, Partha Pratim
中科院分区:
计算机科学4区
文献类型:
--
作者:
Choudhury, Dwaipayan;Xiang, Lizhi;Rajam, Aravind Sukumaran;Kalyanaraman, Ananth;Pande, Partha Pratim

文献摘要

相似文献

图应用程序的工作负载主要是随机内存访问,具有较差的局部性。为了解决计算的不规则性和稀疏性,最近提出了基于ReRAM的内存处理(PIM)架构。这些ReRAM架构设计中的大多数都专注于将图计算映射到一组乘法和累加(MAC)运算中。ReRAM还提供了一个关键优势,即通过允许PIM来减少内核和内存之间的内存延迟。然而,当在基于ReRAM的众核架构上实现时,图形应用程序仍然面临两个关键挑战-显著的存储要求(特别是由于浪费的零单元存储)和大量的片上流量。为了解决这两个挑战,在本文中,我们提出了一种基于3D NoC的ReRAM众核架构的设计。我们提出的架构采用了一种新的交叉感知节点重新排序,以减少ReRAM存储需求。其次,它的3D NoC设计减少了片上通信延迟。我们的架构在基于ReRAM的图形加速方面优于最先进的技术,性能高达5倍,而对于一系列图形输入和工作负载,能耗最多可减少10.3倍。
Graph application workloads are dominated by random memory accesses with the poor locality. To tackle the irregular and sparse nature of computation, ReRAM-based Processing-in-Memory (PIM) architectures have been proposed recently. Most of these ReRAM architecture designs have focused on mapping graph computations into a set of multiply-and-accumulate (MAC) operations. ReRAMs also offer a key advantage in reducing memory latency between cores and memory by allowing for PIM. However, when implemented on a ReRAM-based manycore architecture, graph applications still pose two key challenges—significant storage requirements (particularly due to wasted zero cell storage), and significant amount of on-chip traffic. To tackle these two challenges, in this article, we propose the design of a 3D NoC-enabled ReRAM-based manycore architecture. Our proposed architecture incorporates a novel crossbar-aware node reordering to reduce ReRAM storage requirements. Secondly, its 3D NoC-enabled design reduces on-chip communication latency. Our architecture outperforms the state-of-the-art in ReRAM-based graph acceleration by up to 5× in performance while consuming up to 10.3× less energy for a range of graph inputs and workloads.