Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks

Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks
复制标题

DOI:
10.1109/jssc.2016.2616357
复制
发表时间:
2017-01-01
影响因子:
5.4
通讯作者:
Sze, Vivienne
Sze, Vivienne
中科院分区:
工程技术1区
文献类型:
--
作者:
Chen, Yu-Hsin;Krishna, Tushar;Sze, Vivienne

文献摘要

被引文献

相似文献

Eyeriss是一种用于最先进的深度卷积神经网络(CNN)的加速器。它通过重新配置架构来优化整个系统的能效,包括加速器芯片和片外DRAM,以适应各种CNN形状。CNN在现代人工智能系统中得到了广泛的应用,但也给底层硬件的吞吐量和能效带来了挑战。这是因为它的计算需要大量数据,造成大量数据从芯片上和芯片外移动,这比计算更耗能。因此,将任何CNN形状的数据移动能量成本降至最低是高吞吐量和能源效率的关键。Eyeriss通过在具有168个处理单元的空间架构上使用建议的处理数据流,称为行固定(RS)来实现这些目标。RS数据流重新配置给定形状的计算映射,通过最大限度地在本地重用数据来减少昂贵的数据移动(如DRAM访问),从而优化能效。压缩和数据选通也被应用,以进一步提高能源效率。对于278 mW的AlexNet(批量大小N=4),Eyeriss处理卷积层的速率为35帧/S和0.0029动态随机存取/乘法和累加(MAC);对于VGG-16,在236 mW(N=3)的情况下,Eyeriss处理的卷积层的帧/S和0.0035的动态随机存取/乘法和累加(MAC)。
Eyeriss is an accelerator for state-of-the-art deep convolutional neural networks (CNNs). It optimizes for the energy efficiency of the entire system, including the accelerator chip and off-chip DRAM, for various CNN shapes by reconfiguring the architecture. CNNs are widely used in modern AI systems but also bring challenges on throughput and energy efficiency to the underlying hardware. This is because its computation requires a large amount of data, creating significant data movement from on-chip and off-chip that is more energyconsuming than computation. Minimizing data movement energy cost for any CNN shape, therefore, is the key to high throughput and energy efficiency. Eyeriss achieves these goals by using a proposed processing dataflow, called row stationary (RS), on a spatial architecture with 168 processing elements. RS dataflow reconfigures the computation mapping of a given shape, which optimizes energy efficiency by maximally reusing data locally to reduce expensive data movement, such as DRAM accesses. Compression and data gating are also applied to further improve energy efficiency. Eyeriss processes the convolutional layers at 35 frames/s and 0.0029 DRAM access/multiply and accumulation (MAC) for AlexNet at 278 mW (batch size N = 4), and 0.7 frames/s and 0.0035 DRAM access/MAC for VGG-16 at 236 mW (N = 3).