Intelligence Processing Units Accelerate Neuromorphic Learning

Intelligence Processing Units Accelerate Neuromorphic Learning
复制标题

DOI:
10.48550/arxiv.2211.10725
复制
发表时间:
2022-11
期刊:
ArXiv
影响因子:
--
通讯作者:
P. Sun;A. Titterton;Anjlee Gopiani;Tim Santos;A. Basu;Wei Lu;J. Eshraghian
P. Sun;A. Titterton;Anjlee Gopiani;Tim Santos;A. Basu;Wei Lu;J. Eshraghian
中科院分区:
其他
文献类型:
--
作者:
P. Sun;A. Titterton;Anjlee Gopiani;Tim Santos;A. Basu;Wei Lu;J. Eshraghian

文献摘要

被引文献

相似文献

- 尖峰神经网络(SNN)在使用深度学习工作负载执行推理时,在能耗和延迟方面取得了显著改善。误差反向传播目前被认为是训练SNN的最有效方法,但具有讽刺意味的是,当在现代图形处理单元(GPU)上训练时,这比非尖峰网络更昂贵。Graphcore的智能处理单元(IPU)的出现平衡了深度学习工作负载的并行化性质与训练SNN时普遍存在的操作的顺序,可重用和稀疏艾德性质。IPU通过在较小的数据块上运行单独的处理线程来采用多指令多数据(MIMD)并行性,这对于求解尖峰神经元动态状态方程所需的顺序、非向量化步骤来说是一种自然的适合。我们提出了我们的自定义SNN Python包snnTorch的IPU优化版本,该版本通过利用低级预编译的自定义操作来加速训练SNN工作负载所特有的不规则和稀疏数据访问模式,从而利用细粒度并行性。我们对一套常用的尖峰神经元模型进行了严格的性能评估,并提出了通过半精度训练进一步减少训练运行时间的方法。通过将顺序处理的成本分摊到可向量化的种群代码中,我们最终展示了将特定领域加速器与下一代神经网络集成的潜力。
—Spiking neural networks (SNNs) have achieved or- ders of magnitude improvement in terms of energy consumption and latency when performing inference with deep learning workloads. Error backpropagation is presently regarded as the most effective method for training SNNs, but in a twist of irony, when training on modern graphics processing units (GPUs) this becomes more expensive than non-spiking networks. The emergence of Graphcore’s Intelligence Processing Units (IPUs) balances the parallelized nature of deep learning workloads with the sequential, reusable, and sparsified nature of operations prevalent when training SNNs. IPUs adopt multi-instruction multi-data (MIMD) parallelism by running individual processing threads on smaller data blocks, which is a natural fit for the sequential, non-vectorized steps required to solve spiking neuron dynamical state equations. We present an IPU-optimized release of our custom SNN Python package, snnTorch , which exploits fine-grained parallelism by utilizing low-level, pre-compiled cus- tom operations to accelerate irregular and sparse data access patterns that are characteristic of training SNN workloads. We provide a rigorous performance assessment across a suite of commonly used spiking neuron models, and propose methods to further reduce training run-time via half-precision training. By amortizing the cost of sequential processing into vectorizable population codes, we ultimately demonstrate the potential for integrating domain-specific accelerators with the next generation of neural networks.