Morph-GCNX: A Universal Architecture for High-Performance and Energy-Efficient Graph Convolutional Network Acceleration

Morph-GCNX: A Universal Architecture for High-Performance and Energy-Efficient Graph Convolutional Network Acceleration
复制标题

DOI:
10.1109/tsusc.2023.3313880
复制
发表时间:
2024-03
影响因子:
3.9
通讯作者:
Ke Wang;Hao Zheng;Jiajun Li;A. Louri
Ke Wang;Hao Zheng;Jiajun Li;A. Louri
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ke Wang;Hao Zheng;Jiajun Li;A. Louri

文献摘要

相似文献

尽管当前的图卷积网络(GCN)加速器在众多应用领域取得了显著成功,但这些GCN加速器无法支持各种GCN内部和之间的数据流,也无法适应多样化的GCN应用。在本文中,我们提出了Morph - GCNX,这是一种灵活的GCN加速器架构,用于实现高性能且节能的GCN运算。所提出的设计包含一个灵活的处理单元(PE)阵列,该阵列可在运行时进行划分,以适应单个GCN内不同层或多个并发GCN的计算需求。所提出的Morph - GCNX还包含一种可变形的互连设计,以支持广泛的GCN数据流,并采用各种并行化和数据重用策略来执行GCN。我们还提出了一种硬件 - 应用协同探索技术,该技术对GCN和硬件设计空间进行探索,以确定最佳的PE划分、工作负载分配、数据流和互连配置,目的是提高整体性能和能效。仿真结果表明,与包括HyGCN、AWB - GCN、LW - GCN、GCoD和GCNAX在内的先前设计相比,所提出的Morph - GCNX架构的性能分别提升了18.8倍、2.9倍、1.9倍、1.8倍和2.5倍,DRAM访问次数分别减少了10.8倍、3.7倍、2.2倍、2.5倍和1.3倍,能耗分别降低了13.2倍、5.6倍、2.1倍、2.5倍和1.3倍。
While current Graph Convolutional Networks (GCNs) accelerators have achieved notable success in a wide range of application domains, these GCN accelerators can not support various intra- and inter- GCN dataflows or adapt to diverse GCN applications. In this paper, we propose Morph-GCNX, a flexible GCN accelerator architecture for high-performance and energy-efficient GCN execution. The proposed design consists of a flexible Processing Element (PE) array that can be partitioned at runtime and adapt to the computational needs of different layers within a GCN or multiple concurrent GCNs. The proposed Morph-GCNX also consists of a morphable interconnection design to support a wide range of GCN dataflows with various parallelization and data reuse strategies for GCN execution. We also propose a hardware-application co-exploration technique that explores the GCN and hardware design spaces to identify the best PE partition, workload allocation, dataflow, and interconnection configurations, with the goal of improving overall performance and energy. Simulation results show that the proposed Morph-GCNX architecture achieves 18.8×, 2.9×, 1.9×, 1.8×, and 2.5× better performance, reduces DRAM accesses by a factor of 10.8×, 3.7×, 2.2×, 2.5×, and 1.3×, and improves energy consumption by 13.2×, 5.6×, 2.1×, 2.5×, and 1.3×, as compared to prior designs including HyGCN, AWB-GCN, LW-GCN, GCoD, and GCNAX, respectively.