DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic Hardware

DFSynthesizer: Dataflow-based Synthesis of Spiking Neural Networks to Neuromorphic Hardware
复制标题

DOI:
10.1145/3479156
复制
发表时间:
2021-08
期刊:
ACM Transactions on Embedded Computing Systems (TECS)
影响因子:
--
通讯作者:
Shihao Song;Harry Chong;Adarsha Balaji;Anup Das;J. Shackleford;Nagarajan Kandasamy
Shihao Song;Harry Chong;Adarsha Balaji;Anup Das;J. Shackleford;Nagarajan Kandasamy
中科院分区:
其他
文献类型:
--
作者:
Shihao Song;Harry Chong;Adarsha Balaji;Anup Das;J. Shackleford;Nagarajan Kandasamy

文献摘要

相似文献

尖峰神经网络(SNN)是一种新兴的计算模型,该模型使用事件驱动的激活和生物启发的学习算法。基于SNN的机器学习程序通常在基于瓷砖的神经形态硬件平台上执行,每个瓷砖都由一个称为横杆的计算单元组成,该计算单元绘制了该程序的神经元和突触。但是,在现成的神经形态硬件上合成此类程序是具有挑战性的。这是因为硬件的固有资源和延迟限制,这既影响模型性能,例如准确性和硬件性能,例如吞吐量。我们提出了DFSynthesizer,这是将基于SNN的机器学习程序合成神经形态硬件的端到端框架。提出的框架分为四个步骤。首先,它分析机器学习程序并使用代表性数据生成SNN工作负载。其次,它可以分区SNN的工作负载,并生成适合目标神经形态硬件横梁的簇。第三,它利用同步数据流图(SDFG)的丰富语义代表一个群集的SNN程序,可以根据关键硬件约束(例如跨键数,每个横梁的维度,瓷砖上的缓冲区和瓷砖)进行性能分析通信带宽。最后,它使用一种新颖的调度算法来在硬件的横梁上执行簇,从而保证了硬件性能。我们通过10种常用的机器学习程序评估DFSynthesizer。我们的结果表明,与当前的映射方法相比,DFSynthesizer提供了更严格的性能保证。
Spiking Neural Networks (SNNs) are an emerging computation model that uses event-driven activation and bio-inspired learning algorithms. SNN-based machine learning programs are typically executed on tile-based neuromorphic hardware platforms, where each tile consists of a computation unit called a crossbar, which maps neurons and synapses of the program. However, synthesizing such programs on an off-the-shelf neuromorphic hardware is challenging. This is because of the inherent resource and latency limitations of the hardware, which impact both model performance, e.g., accuracy, and hardware performance, e.g., throughput. We propose DFSynthesizer, an end-to-end framework for synthesizing SNN-based machine learning programs to neuromorphic hardware. The proposed framework works in four steps. First, it analyzes a machine learning program and generates SNN workload using representative data. Second, it partitions the SNN workload and generates clusters that fit on crossbars of the target neuromorphic hardware. Third, it exploits the rich semantics of the Synchronous Dataflow Graph (SDFG) to represent a clustered SNN program, allowing for performance analysis in terms of key hardware constraints such as number of crossbars, dimension of each crossbar, buffer space on tiles, and tile communication bandwidth. Finally, it uses a novel scheduling algorithm to execute clusters on crossbars of the hardware, guaranteeing hardware performance. We evaluate DFSynthesizer with 10 commonly used machine learning programs. Our results demonstrate that DFSynthesizer provides a much tighter performance guarantee compared to current mapping approaches.