DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning

DianNao: A Small-Footprint High-Throughput Accelerator for Ubiquitous Machine-Learning
复制标题

DOI:
10.1145/2541940.2541967
复制
发表时间:
2014-04-01
影响因子:
--
通讯作者:
Temam, Olivier
Temam, Olivier
中科院分区:
其他
文献类型:
--
作者:
Chen, Tianshi;Du, Zidong;Temam, Olivier

文献摘要

被引文献

相似文献

机器学习任务在广泛的领域和广泛的系统(从嵌入式系统到数据中心)中变得无处不在。与此同时,一小部分机器学习算法(尤其是卷积和深度神经网络,即cnn和dnn)在许多应用中被证明是最先进的。随着架构向由核心和加速器混合组成的异构多核发展,机器学习加速器可以实现罕见的效率(由于目标算法数量少)和广泛应用范围的结合。到目前为止,大多数机器学习加速器的设计都集中在有效地实现算法的计算部分。然而,最近最先进的cnn和dnn的特点是它们的大尺寸。在本研究中,我们设计了一个用于大规模cnn和dnn的加速器,特别强调了内存对加速器设计、性能和能量的影响。我们表明,可以设计一个具有高吞吐量的加速器,能够在3.02 mm(2)和485 mW的小占地中执行452 GOP/s(关键的神经网络操作,如突触权重乘法和神经元输出加法);与128位2GHz SIMD处理器相比,加速器速度提高了117.87倍,总能耗降低21.08倍。在65nm布置后得到了加速器的特性。在很小的空间内实现如此高的吞吐量,可以在广泛的系统和广泛的应用中使用最先进的机器学习算法。
Machine-Learning tasks are becoming pervasive in a broad range of domains, and in a broad range of systems (from embedded systems to data centers). At the same time, a small set of machine-learning algorithms (especially Convolutional and Deep Neural Networks, i.e., CNNs and DNNs) are proving to be state-of-the-art across many applications. As architectures evolve towards heterogeneous multi-cores composed of a mix of cores and accelerators, a machine-learning accelerator can achieve the rare combination of efficiency (due to the small number of target algorithms) and broad application scope.Until now, most machine-learning accelerator designs have focused on efficiently implementing the computational part of the algorithms. However, recent state-of-the-art CNNs and DNNs are characterized by their large size. In this study, we design an accelerator for large-scale CNNs and DNNs, with a special emphasis on the impact of memory on accelerator design, performance and energy.We show that it is possible to design an accelerator with a high throughput, capable of performing 452 GOP/s (key NN operations such as synaptic weight multiplications and neurons outputs additions) in a small footprint of 3.02 mm(2) and 485 mW; compared to a 128-bit 2GHz SIMD processor, the accelerator is 117.87x faster, and it can reduce the total energy by 21.08x. The accelerator characteristics are obtained after layout at 65nm. Such a high throughput in a small footprint can open up the usage of state-of-the-art machine-learning algorithms in a broad set of systems and for a broad set of applications.