Efficient Spectral Graph Convolutional Network Deployment on Memristive Crossbars

Efficient Spectral Graph Convolutional Network Deployment on Memristive Crossbars
复制标题

DOI:
10.1109/tetci.2022.3210998
复制
发表时间:
2023-04
影响因子:
5.3
通讯作者:
Bo Lyu;Maher Hamdi;Yin Yang;Yuting Cao;Zheng Yan;Ke Li;Shiping Wen;Tingwen Huang
Bo Lyu;Maher Hamdi;Yin Yang;Yuting Cao;Zheng Yan;Ke Li;Shiping Wen;Tingwen Huang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Bo Lyu;Maher Hamdi;Yin Yang;Yuting Cao;Zheng Yan;Ke Li;Shiping Wen;Tingwen Huang

文献摘要

被引文献

相似文献

图神经网络(GNNs)由于其对图结构知识建模的卓越能力而引起了越来越多的研究兴趣。然而,GNN遭受密集的数据交换和较差的数据局部性,这将在传统的基于互补金属氧化物半导体(CMOS)的von-Neumann计算架构(图形处理单元(GPU)、中央处理单元(CPU))下由于“存储墙”问题而导致关键的性能和能量瓶颈。幸运的是,基于忆阻交叉杆的计算已经成为最有前途的神经形态计算架构之一,它已被广泛研究作为卷积神经网络(CNN),递归神经网络(RNN),尖峰神经网络(SNNs)等的计算平台。进一步,基于GCN(极高稀疏性和非零数据不平衡分布)的结构和忆阻交叉开关电路的神经形态特性,提出了稀疏拉普拉斯矩阵重排和对角块矩阵乘法相结合的加速方法。在监督学习图数据集(QM 7)上,忆阻器交叉杆的模拟实验达到了90.3%的总体准确率,与原始计算相比,本文提出的加速计算框架(对角块大小减半)实现了忆阻器数量减少27.3%。此外,在无监督学习数据集(空手道俱乐部)上,我们的方法在使用半尺寸对角块映射时没有损失准确性,并且忆阻器数量减少了32.2%。
Graph Neural Networks (GNNs) have attracted increasing research interest for their remarkable capability to model graph-structured knowledge. However, GNNs suffer from intensive data exchange and poor data locality, which will cause critical performance and energy bottlenecks under conventional complementary metal oxide semiconductor (CMOS)-based von-Neumann computing architectures (graphics processing unit (GPU), central processing unit (CPU)) for the “Memory Wall” issue. Fortunately, memristive crossbar-based computation has emerged as one of the most promising neuromorphic computing architectures, which has been widely studied as the computing platform for convolutional neural network (CNNs), recurrent neural network (RNNs), spiking neural network (SNNs), etc. This paper proposes the deployment of spectral graph convolutional networks (GCNs) on memristive crossbars. Further, based on the structure of GCNs (extremely high sparsity and unbalanced non-zero data distribution) and the neuromorphic characteristics of memristive crossbar circuit, we propose the acceleration method that consists of Sparse Laplace Matrix Reordering and Diagonal Block Matrix Multiplication. The simulated experiment on memristor crossbars achieves 90.3% overall accuracy on the supervised learning graph dataset (QM7), and compared with the original computation, the proposed acceleration computing framework (with half-size diagonal blocks) achieves a 27.3% reduction of memristor number. Additionally, on the unsupervised learning dataset (Karate club), our method shows no loss of accuracy with half-size diagonal block mapping and reaches a 32.2% reduction of memristor number.