Chain-NN: An energy-efficient 1D chain architecture for accelerating deep convolutional neural networks

Chain-NN: An energy-efficient 1D chain architecture for accelerating deep convolutional neural networks
复制标题

DOI:
10.23919/date.2017.7927142
复制
发表时间:
2017-03
期刊:
Design, Automation & Test in Europe Conference & Exhibition (DATE), 2017
影响因子:
--
通讯作者:
Shihao Wang;Dajiang Zhou;Xushen Han;T. Yoshimura
Shihao Wang;Dajiang Zhou;Xushen Han;T. Yoshimura
中科院分区:
其他
文献类型:
--
作者:
Shihao Wang;Dajiang Zhou;Xushen Han;T. Yoshimura

文献摘要

相似文献

深卷积神经网络(CNN)在许多计算机视觉任务中表现出了良好的性能。然而,CNN的高计算复杂性涉及到计算处理器内核和存储层次之间的海量数据移动,这占据了功耗的主要部分。提出了一种新的能量高效的一维链结构--Chain-NN,用于加速深层CNN。Chain-NN由专用的双通道过程引擎(PE)组成。在Chain-NN中,卷积是由一组相邻PE组成的一维脉动基元完成的。这些脉动基元与所提出的逐列扫描输入模式相结合,可以充分重用输入操作数,从而降低节能所需的存储带宽。此外,一维链结构允许根据特定的CNN参数轻松地重新配置脉动基元,而设计复杂度更低。链状神经网络的合成和布局采用台积电28 nm工艺。它的成本为3751K逻辑门和352KB的片内存储器。结果表明,576个PE链神经网络可以扩展到700 MHz。这在567.5 mW的情况下实现了806.4GOPS的峰值吞吐量,并且能够以326.2fps的帧速率加速AlexNet中的五个卷积层。1421.0GOPS/W的功率效率至少是最先进作品的2.5到4.1倍。
Deep convolutional neural networks (CNN) have shown their good performances in many computer vision tasks. However, the high computational complexity of CNN involves a huge amount of data movements between the computational processor core and memory hierarchy which occupies the major of the power consumption. This paper presents Chain-NN, a novel energy-efficient 1D chain architecture for accelerating deep CNNs. Chain-NN consists of the dedicated dual-channel process engines (PE). In Chain-NN, convolutions are done by the 1D systolic primitives composed of a group of adjacent PEs. These systolic primitives, together with the proposed column-wise scan input pattern, can fully reuse input operand to reduce the memory bandwidth requirement for energy saving. Moreover, the 1D chain architecture allows the systolic primitives to be easily reconfigured according to specific CNN parameters with fewer design complexity. The synthesis and layout of Chain-NN is under TSMC 28nm process. It costs 3751k logic gates and 352KB on-chip memory. The results show a 576-PE Chain-NN can be scaled up to 700MHz. This achieves a peak throughput of 806.4GOPS with 567.5mW and is able to accelerate the five convolutional layers in AlexNet at a frame rate of 326.2fps. 1421.0GOPS/W power efficiency is at least 2.5 to 4.1x times better than the state-of-the-art works.