RAPIDNN: In-Memory Deep Neural Network Acceleration Framework

RAPIDNN: In-Memory Deep Neural Network Acceleration Framework
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
通讯作者:
M. Imani;Mohammad Samragh;Yeseong Kim;Saransh Gupta;F. Koushanfar;Tajana Simunic
M. Imani;Mohammad Samragh;Yeseong Kim;Saransh Gupta;F. Koushanfar;Tajana Simunic
中科院分区:
其他
文献类型:
--
作者:
M. Imani;Mohammad Samragh;Yeseong Kim;Saransh Gupta;F. Koushanfar;Tajana Simunic

文献摘要

被引文献

相似文献

深度神经网络(DNN)在图像处理、视频分割、语音识别等领域有着广泛的应用前景。在当前系统上运行最先进的DNN主要依赖于通用处理器、ASIC设计或FPGA加速器,由于片上存储器和数据传输带宽有限,所有这些都受到数据移动的影响。在这项工作中,我们提出了一个新的框架,称为RAPIDNN,它在内存中处理所有的DNN操作,以最大限度地减少数据移动的成本。为了实现内存中处理,RAPIDNN重新解释DNN模型并将其映射到专用加速器中,该加速器使用非易失性存储块设计,对四种基本DNN操作进行建模,即乘法、加法、激活函数和池化。该框架使用聚类方法提取DNN模型的代表性操作数,例如权重和输入值,以优化该模型以用于存储器内处理。然后,它将提取的操作数及其预计算结果映射到加速器存储块中。在运行时,加速器基于高效的内存搜索能力来识别计算结果,这也提供了近似值的可调性,以进一步提高计算效率。我们的评估显示,与最先进的DNN加速器Isaac和Pipelayer相比,RAPIDNN的能效分别提高了68.4倍和49.5倍,加速比分别提高了48.1倍和10.9倍,同时保证了不到0.3%的质量损失。
Deep neural networks (DNN) have demonstrated effectiveness for various applications such as image processing, video segmentation, and speech recognition. Running state-of-the-art DNNs on current systems mostly relies on either generalpurpose processors, ASIC designs, or FPGA accelerators, all of which suffer from data movements due to the limited onchip memory and data transfer bandwidth. In this work, we propose a novel framework, called RAPIDNN, which processes all DNN operations within the memory to minimize the cost of data movement. To enable in-memory processing, RAPIDNN reinterprets a DNN model and maps it into a specialized accelerator, which is designed using non-volatile memory blocks that model four fundamental DNN operations, i.e., multiplication, addition, activation functions, and pooling. The framework extracts representative operands of a DNN model, e.g., weights and input values, using clustering methods to optimize the model for in-memory processing. Then, it maps the extracted operands and their precomputed results into the accelerator memory blocks. At runtime, the accelerator identifies computation results based on efficient in-memory search capability which also provides tunability of approximation to further improve computation efficiency. Our evaluation shows that RAPIDNN achieves 68.4x, 49.5x energy efficiency improvement and 48.1x, 10.9x speedup as compared to ISAAC and PipeLayer, the state-of-the-art DNN accelerators, while ensuring less than 0.3% of quality loss.