GPUs Outperform Current HPC and Neuromorphic Solutions in Terms of Speed and Energy When Simulating a Highly-Connected Cortical Model

GPUs Outperform Current HPC and Neuromorphic Solutions in Terms of Speed and Energy When Simulating a Highly-Connected Cortical Model
复制标题

DOI:
10.3389/fnins.2018.00941
复制
发表时间:
2018-12-12
影响因子:
4.3
通讯作者:
Nowotny, Thomas
Nowotny, Thomas
中科院分区:
医学2区
文献类型:
--
作者:
Knight, James C.;Nowotny, Thomas

文献摘要

被引文献

相似文献

虽然神经形态系统可能是部署尖峰神经网络 (SNN) 的终极平台,但它们的分布式特性和针对特定类型模型的优化使得它们成为开发这些模型的笨拙工具。相反,SNN 模型往往在具有标准冯·诺依曼 CPU 架构的计算机或计算机集群上开发和模拟。在过去的十年中,NVIDIA GPU 加速器不仅成为许多工作站的常见设备,而且还进入了高性能计算领域,目前已在全球 10 强超级计算站点中的 50% 中使用。在本文中,我们使用 GeNN 代码生成器在 GPU 硬件上重新实现两个受新皮层启发的电路规模点神经元网络模型。我们根据在传统 HPC 硬件上运行 NEST 获得的先前结果验证 GPU 模拟的正确性,并将速度和能耗方面的性能与基于 CPU 的 HPC 和神经形态硬件的已发布数据进行比较。使用单个 NVIDIA Tesla V100 加速器可以以接近 0.5 倍实时的速度模拟皮质柱的全尺寸模型,比目前使用基于 CPU 的集群或 SpiNNaker 神经拟态系统的速度更快。此外,我们发现,在一系列 GPU 系统中,微电路模拟的求解能量以及每个突触事件的能量比 SpiNNaker 或基于 CPU 的模拟低 14 倍。除了模拟速度和能耗方面的性能之外,模型的有效初始化也是一个至关重要的问题,特别是在需要重复运行和参数空间探索的研究环境中。因此,我们还在本文中介绍了最新版本 GeNN 中实现的一些新颖的并行初始化方法,并演示了它们如何实现进一步的速度和能量优势。
While neuromorphic systems may be the ultimate platform for deploying spiking neural networks (SNNs), their distributed nature and optimization for specific types of models makes them unwieldy tools for developing them. Instead, SNN models tend to be developed and simulated on computers or clusters of computers with standard von Neumann CPU architectures. Over the last decade, as well as becoming a common fixture in many workstations, NVIDIA GPU accelerators have entered the High Performance Computing field and are now used in 50 % of the Top 10 super computing sites worldwide. In this paper we use our GeNN code generator to re-implement two neo-cortex-inspired, circuit-scale, point neuron network models on GPU hardware. We verify the correctness of our GPU simulations against prior results obtained with NEST running on traditional HPC hardware and compare the performance with respect to speed and energy consumption against published data from CPU-based HPC and neuromorphic hardware. A full-scale model of a cortical column can be simulated at speeds approaching 0.5 x real-time using a single NVIDIA Tesla V100 accelerator-faster than is currently possible using a CPU based cluster or the SpiNNaker neuromorphic system. In addition, we find that, across a range of GPU systems, the energy to solution as well as the energy per synaptic event of the microcircuit simulation is as much as 14 x lower than either on SpiNNaker or in CPU-based simulations. Besides performance in terms of speed and energy consumption of the simulation, efficient initialization of models is also a crucial concern, particularly in a research context where repeated runs and parameter-space exploration are required. Therefore, we also introduce in this paper some of the novel parallel initialization methods implemented in the latest version of GeNN and demonstrate how they can enable further speed and energy advantages.