Reconfigurable Dataflow Optimization for Spatiotemporal Spiking Neural Computation on Systolic Array Accelerators

Reconfigurable Dataflow Optimization for Spatiotemporal Spiking Neural Computation on Systolic Array Accelerators
复制标题

DOI:
10.1109/iccd50377.2020.00027
复制
发表时间:
2020-10
期刊:
2020 IEEE 38th International Conference on Computer Design (ICCD)
影响因子:
--
通讯作者:
Jeong-Jun Lee;Peng Li
Jeong-Jun Lee;Peng Li
中科院分区:
其他
文献类型:
--
作者:
Jeong-Jun Lee;Peng Li

文献摘要

相似文献

尖峰神经网络(SNN)提供了一个有前途的生物合理的计算模型,并借给自己的超低功耗事件驱动处理的神经形态处理器。与传统的人工神经网络相比,SNN非常适合处理复杂的时空数据。尽管它的意义,尖峰神经加速器架构的低优化尚未得到广泛的研究。认识到需要有效地处理复杂的时空数据,同时考虑到尖峰活动的全或无性质,我们提出了整体可重构的卷积神经网络(S-CNN)脉动阵列加速优化。介绍了一种新的方案,用于跨多个时间点的并行加速计算,其进一步允许系统优化可变平铺以获得大的性能和效率增益。我们展示了如何可变平铺,特别是时间维度的定位,可以有针对性地优化数据移动,吞吐量和能源效率。此外,我们还探索了联合层相关的高速缓存和加速器硬件优化,以进一步提高性能和能效。为了支持系统性的设计空间探索,我们开发了一个SNN低功耗模拟器,能够分析任何目标S-CNN的脉动阵列加速器的吞吐量和能量耗散,同时考虑尖峰神经计算的固有时空特性。所提出的技术在吞吐量、能效和延迟能量积方面提供了数量级的改进,以加速深度Alexnet和VGG-16 SNN。
Spiking neural networks (SNNs) offer a promising biologically-plausible computing model and lend themselves to ultra-low-power event-driven processing on neuromorphic processors. Compared with the conventional artificial neural networks, SNNs are well-suited for processing complex spatiotemporal data. Despite its significance, dataflow optimization of spiking neural accelerator architectures has not been extensively studied. Recognizing the need for efficient processing of complex spatiotemporal data while considering the all-or-none nature of spiking activities, we propose holistic reconfigurable dataflow optimization for systolic array acceleration of spiking convolutional networks (S-CNNs). A novel scheme is introduced for parallel acceleration of computation across multiple time points, which further allows for systemic optimization of variable tiling for a large performance and efficiency gains. We show how variable tiling, in particular, the positioning of the temporal dimension, can be targeted to optimize data movement, throughput, and energy efficiency. Furthermore, we explore joint layer-dependent dataflow and accelerator hardware optimization to further boost performance and energy efficiency. To support systemic design space exploration, we develop an SNN dataflow simulator capable of analyzing the throughput and energy dissipation of systolic array accelerators for any targeted S-CNN while considering the inherent spatiotemporal characteristics of spiking neural computation. The proposed techniques deliver orders of magnitude of improvements on throughput, energy efficiency, and delay-energy product for accelerating deep Alexnet and VGG-16 SNNs.