Dataflow-Based Mapping of Spiking Neural Networks on Neuromorphic Hardware

Dataflow-Based Mapping of Spiking Neural Networks on Neuromorphic Hardware
复制标题

DOI:
10.1145/3194554.3194627
复制
发表时间:
2018-05
期刊:
Proceedings of the 2018 Great Lakes Symposium on VLSI
影响因子:
--
通讯作者:
Anup Das;Akash Kumar
Anup Das;Akash Kumar
中科院分区:
其他
文献类型:
--
作者:
Anup Das;Akash Kumar

文献摘要

被引文献

相似文献

尖峰神经网络(SNN)是用于模式识别和图像分类应用程序的强大计算引擎。除了识别和分类精度等应用程序性能外,在硬件上执行这些应用程序时,系统性能(例如吞吐量)也变得很重要。我们建议在基于横杆的神经形态硬件上绘制基于SNN的应用程序的系统设计流,从而确保应用程序和系统性能。同步数据流图(SDFGS)用于用扩展语义来建模这些应用程序,以表示神经网络拓扑。然后,使用自定义的调度来分析吞吐量,并结合了硬件约束,例如突触记忆,通信和横梁的I/O带宽。我们的Design-Flow集成了与SDF3的GPU加速应用程序级模拟器Carlsim,这是用于在硬件上映射SDFG的工具。我们对代表性的神经形态硬件进行了实验和合成SNN的实验,展示了给定应用程序性能的吞吐量 - 资源折衷。对于吞吐量受限的应用程序,我们平均显示了20%的硬件使用量,能源消耗降低了19%。对于可吞吐量的应用程序,与最先进的方法相比,我们的吞吐量平均高出53%。
Spiking Neural Networks (SNNs) are powerful computation engines for pattern recognition and image classification applications. Apart from application performance such as recognition and classification accuracy, system performance such as throughput becomes important when executing these applications on a hardware. We propose a systematic design-flow to map SNN-based applications on a crossbar-based neuromorphic hardware, guaranteeing application as well as system performance. Synchronous Dataflow Graphs (SDFGs) are used to model these applications with extended semantics to represent neural network topologies. Self-timed scheduling is then used to analyze throughput, incorporating hardware constraints such as synaptic memory, communication and I/O bandwidth of crossbars. Our design-flow integrates CARLsim, a GPU-accelerated application-level SNN simulator with SDF3, a tool for mapping SDFG on hardware. We conducted experiments with realistic and synthetic SNNs on representative neuromorphic hardware, demonstrating throughput-resource trade-offs for a given application performance. For throughput-constrained applications, we show average 20% reduction of hardware usage with 19% reduction in energy consumption. For throughput-scalable applications, we show an average 53% higher throughput compared to a state-of-the-art approach.