GCNAX: A Flexible and Energy-efficient Accelerator for Graph Convolutional Neural Networks

GCNAX: A Flexible and Energy-efficient Accelerator for Graph Convolutional Neural Networks
复制标题

DOI:
10.1109/hpca51647.2021.00070
复制
发表时间:
2021-02
期刊:
2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Jiajun Li;A. Louri;Avinash Karanth;Razvan C. Bunescu
Jiajun Li;A. Louri;Avinash Karanth;Razvan C. Bunescu
中科院分区:
其他
文献类型:
--
作者:
Jiajun Li;A. Louri;Avinash Karanth;Razvan C. Bunescu

文献摘要

被引文献

相似文献

图卷积神经网络(GCN)已成为扩展深度学习用于图数据分析的有效方法。鉴于图通常是不规则的,因为图中的节点可能具有不同数量的邻居,因此有效地处理GCN对底层硬件构成了重大挑战。尽管已经提出了专用GCN加速器以提供优于通用处理器的性能,但是现有加速器不仅未充分利用计算引擎,而且还强加了冗余数据访问,这降低了吞吐量和能量效率。因此,优化计算引擎和存储器之间的总体数据流,即,GCN调度流是实现GCN高效处理的关键,它能最大限度地提高资源利用率,最小限度地减少数据移动。2在本文中,我们提出了一种灵活的GCN调度流,它能同时提高资源利用率和减少数据移动。这是通过充分探索GCN的设计空间和评估的执行周期和DRAM访问的数量通过分析框架。与传统的GCN环路采用严格的环路阶数和环路融合策略不同,该环路可以重新配置环路阶数和环路融合策略,以适应不同的GCN配置,从而大大提高了效率。然后,我们介绍了一种新的加速器架构,称为GCNAX,它量身定制的计算引擎,缓冲区结构和大小的基础上提出的算法。在五个真实世界的图形数据集上进行评估,我们的模拟结果表明,GCNAX将DRAM访问减少了8.1倍和2.4倍,同时实现了8.9倍,1.6倍的加速比和9.5倍,2.3倍的能源节省,分别超过HyGCN和AWB-GCN。
Graph convolutional neural networks (GCNs) have emerged as an effective approach to extend deep learning for graph data analytics. Given that graphs are usually irregular, as nodes in a graph may have a varying number of neighbors, processing GCNs efficiently pose a significant challenge on the underlying hardware. Although specialized GCN accelerators have been proposed to deliver better performance over generic processors, prior accelerators not only under-utilize the compute engine, but also impose redundant data accesses that reduce throughput and energy efficiency. Therefore, optimizing the overall flow of data between compute engines and memory, i.e., the GCN dataflow, which maximizes utilization and minimizes data movement is crucial for achieving efficient GCN processing.In this paper, we propose a flexible and optimized dataflow for GCNs that simultaneously improves resource utilization and reduces data movement. This is realized by fully exploring the design space of GCN dataflows and evaluating the number of execution cycles and DRAM accesses through an analysis framework. Unlike prior GCN dataflows, which employ rigid loop orders and loop fusion strategies, the proposed dataflow can reconFigure the loop order and loop fusion strategy to adapt to different GCN configurations, which results in much improved efficiency. We then introduce a novel accelerator architecture called GCNAX, which tailors the compute engine, buffer structure and size based on the proposed dataflow. Evaluated on five real-world graph datasets, our simulation results show that GCNAX reduces DRAM accesses by a factor of $8.1 \times$ and $2.4 \times$, while achieving $8.9 \times, 1.6 \times$ speedup and $9.5 \times$, $2.3 \times$ energy savings on average over HyGCN and AWB-GCN, respectively.