Acceleration of Deep Neural Network Training with Resistive Cross-Point Devices: Design Considerations.

Acceleration of Deep Neural Network Training with Resistive Cross-Point Devices: Design Considerations.
复制标题

使用电阻交叉点设备的深度神经网络训练的加速:设计注意事项。

DOI:
10.3389/fnins.2016.00333
复制
发表时间:
2016
影响因子:
4.3
通讯作者:
Vlasov Y
Vlasov Y
中科院分区:
医学2区
文献类型:
--
作者:
Gokmen T;Vlasov Y

文献摘要

被引文献

相似文献

近年来,深度神经网络(DNN)在语音识别、视觉目标检测、模式提取等大规模分析和分类任务中显示出了巨大的商业影响力。然而,人们普遍认为大型DNN的训练是一项耗时和计算密集型的任务,需要花费数天的数据中心级计算资源。在这里,我们提出了一种阻性处理单元(RPU)的概念,它可以在使用更少的功率的情况下,潜在地将DNN的训练速度加快数量级。所提出的RPU设备可以在本地存储和更新权值,从而最小化训练过程中的数据移动,并允许充分利用训练算法的局部性和并行性。我们评估了各种RPU器件特性/非理想性和系统参数对性能的影响,以便推导出用于DNN培训的加速器芯片的器件和系统级规范,以实现现实的CMOS兼容技术。对于具有约10亿权重的大型DNN,这种大规模并行RPU架构与最先进的微处理器相比可以实现30,000倍的加速系数,同时提供84,000GigaOps/S/W的能效。目前需要在数据中心大小的包含数千台机器的集群上进行数天培训的问题,可以在单个RPU加速器上在数小时内得到解决。由RPU加速器集群组成的系统将能够处理具有数万亿参数的大数据问题,这些问题在今天是不可能解决的,例如,自然语音识别和所有世界语言之间的翻译,对大量商业和科学数据流的实时分析,以及对来自大量物联网(IoT)传感器的多模式感觉数据流的集成和分析。
In recent years, deep neural networks (DNN) have demonstrated significant business impact in large scale analysis and classification tasks such as speech recognition, visual object detection, pattern extraction, etc. Training of large DNNs, however, is universally considered as time consuming and computationally intensive task that demands datacenter-scale computational resources recruited for many days. Here we propose a concept of resistive processing unit (RPU) devices that can potentially accelerate DNN training by orders of magnitude while using much less power. The proposed RPU device can store and update the weight values locally thus minimizing data movement during training and allowing to fully exploit the locality and the parallelism of the training algorithm. We evaluate the effect of various RPU device features/non-idealities and system parameters on performance in order to derive the device and system level specifications for implementation of an accelerator chip for DNN training in a realistic CMOS-compatible technology. For large DNNs with about 1 billion weights this massively parallel RPU architecture can achieve acceleration factors of 30, 000 × compared to state-of-the-art microprocessors while providing power efficiency of 84, 000 GigaOps∕s∕W. Problems that currently require days of training on a datacenter-size cluster with thousands of machines can be addressed within hours on a single RPU accelerator. A system consisting of a cluster of RPU accelerators will be able to tackle Big Data problems with trillions of parameters that is impossible to address today like, for example, natural speech recognition and translation between all world languages, real-time analytics on large streams of business and scientific data, integration, and analysis of multimodal sensory data flows from a massive number of IoT (Internet of Things) sensors.