Fused-layer CNN accelerators

Fused-layer CNN accelerators
复制标题

DOI:
10.1109/micro.2016.7783725
复制
发表时间:
2016-10
期刊:
2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Manoj Alwani;Han Chen;M. Ferdman;Peter Milder
Manoj Alwani;Han Chen;M. Ferdman;Peter Milder
中科院分区:
其他
文献类型:
--
作者:
Manoj Alwani;Han Chen;M. Ferdman;Peter Milder

文献摘要

被引文献

相似文献

深度卷积神经网络 (CNN) 正在迅速成为计算机视觉的主导方法,也是许多其他普遍机器学习任务的主要组成部分,例如语音识别、自然语言处理和欺诈检测。因此,用于高效评估 CNN 的加速器正在迅速普及。设计此类 CNN 加速器的传统方法是专注于创建加速器来迭代处理 CNN 层。然而,通过处理每一层直至完成,加速器设计必须使用片外存储器来存储层之间的中间数据,因为中间数据太大而无法容纳在芯片上。在这项工作中,我们观察到 CNN 加速器的设计空间中存在一个先前未探索的维度,该维度专注于跨卷积层的数据流。我们发现,我们能够通过修改输入数据进入芯片的顺序来融合多个 CNN 层的处理,从而能够缓存相邻 CNN 层评估之间的中间数据。我们通过为 VGGNet-E 网络的前五个卷积层构建融合层 CNN 加速器,并将其与 Xilinx Virtex-7 FPGA 上实现的最先进的加速器进行比较,证明了我们方法的有效性。我们发现,通过使用 362KB 片上存储,我们的融合层加速器最大限度地减少了片外特征图数据传输,将总传输量减少了 95%,从每个图像 77MB 减少到 3.6MB。
Deep convolutional neural networks (CNNs) are rapidly becoming the dominant approach to computer vision and a major component of many other pervasive machine learning tasks, such as speech recognition, natural language processing, and fraud detection. As a result, accelerators for efficiently evaluating CNNs are rapidly growing in popularity. The conventional approaches to designing such CNN accelerators is to focus on creating accelerators to iteratively process the CNN layers. However, by processing each layer to completion, the accelerator designs must use off-chip memory to store intermediate data between layers, because the intermediate data are too large to fit on chip. In this work, we observe that a previously unexplored dimension exists in the design space of CNN accelerators that focuses on the dataflow across convolutional layers. We find that we are able to fuse the processing of multiple CNN layers by modifying the order in which the input data are brought on chip, enabling caching of intermediate data between the evaluation of adjacent CNN layers. We demonstrate the effectiveness of our approach by constructing a fused-layer CNN accelerator for the first five convolutional layers of the VGGNet-E network and comparing it to the state-of-the-art accelerator implemented on a Xilinx Virtex-7 FPGA. We find that, by using 362KB of on-chip storage, our fused-layer accelerator minimizes off-chip feature map data transfer, reducing the total transfer by 95%, from 77MB down to 3.6MB per image.