Software/Hardware Co-design of 3D NoC-based GPU Architectures for Accelerated Graph Computations

Software/Hardware Co-design of 3D NoC-based GPU Architectures for Accelerated Graph Computations
复制标题

DOI:
10.1145/3514354
复制
发表时间:
2022-04
期刊:
ACM Transactions on Design Automation of Electronic Systems (TODAES)
影响因子:
--
通讯作者:
Dwaipayan Choudhury;Reet Barik;Aravind Sukumaran Rajam;A. Kalyanaraman;And Partha Pratim Pande
Dwaipayan Choudhury;Reet Barik;Aravind Sukumaran Rajam;A. Kalyanaraman;And Partha Pratim Pande
中科院分区:
其他
文献类型:
--
作者:
Dwaipayan Choudhury;Reet Barik;Aravind Sukumaran Rajam;A. Kalyanaraman;And Partha Pratim Pande

文献摘要

相似文献

Manycore图形处理器架构已经成为加速图形计算的中流砥柱。在许多核心架构上,图计算性能的主要瓶颈之一是数据移动。由于图处理中的大多数访问都是通过顶点邻域查找进行的,因此图数据结构中的局部性对数据移动的程度起着关键作用。顶点重排序是一种广泛使用的技术,用于改善图数据结构中的数据局部性。然而,这些重新排序方案本身是不够的,因为它们需要在多核GPU架构上用有效的任务分配来补充,以减少由于本地高速缓存未命中而导致的延迟。因此,在本文中,我们介绍了一种用于加速图计算的软硬件协同设计框架。我们的方法结合了体系结构感知的顶点重新排序和基于优先级的任务分配技术。由于任务分配的目的是减少片上时延和相关能量,因此在多核平台中选择片上网络(NoC)作为通信骨干是一个重要的参数。通过利用新兴的三维(3D)集成技术,我们提出了一种支持小世界片上网络(SWNoC)的多核GPU架构的设计,其中连接流多处理器(SM)和存储控制器(MC)的链路的布局遵循幂函数分布。与传统平面网格体系结构上运行的数据集的默认顺序相比,所提出的支持3D SWNoC的软硬件协同设计框架实现了11.1%到22.9%的性能提升和16.4%到32.6%的能耗降低,具体取决于数据集和图形应用程序。
Manycore GPU architectures have become the mainstay for accelerating graph computations. One of the primary bottlenecks to performance of graph computations on manycore architectures is the data movement. Since most of the accesses in graph processing are due to vertex neighborhood lookups, locality in graph data structures plays a key role in dictating the degree of data movement. Vertex reordering is a widely used technique to improve data locality within graph data structures. However, these reordering schemes alone are not sufficient as they need to be complemented with efficient task allocation on manycore GPU architectures to reduce latency due to local cache misses. Consequently, in this article, we introduce a software/hardware co-design framework for accelerating graph computations. Our approach couples an architecture-aware vertex reordering with a priority-based task allocation technique. As the task allocation aims to reduce on-chip latency and associated energy, the choice of Network-on-Chip (NoC) as the communication backbone in the manycore platform is an important parameter. By leveraging emerging three-dimensional (3D) integration technology, we propose design of a small-world NoC (SWNoC)-enabled manycore GPU architecture, where the placement of the links connecting the streaming multiprocessors (SMs) and the memory controllers (MCs) follow a power-law distribution. The proposed 3D SWNoC-enabled software/hardware co-design framework achieves 11.1% to 22.9% performance improvement and 16.4% to 32.6% less energy consumption depending on the dataset and the graph application, when compared to the default order of dataset running on a conventional planar mesh architecture.