Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework

Towards Ultra-High Performance and Energy Efficiency of Deep Learning Systems: An Algorithm-Hardware Co-Optimization Framework
复制标题

DOI:
10.1609/aaai.v32i1.11653
复制
发表时间:
2018-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Yanzhi Wang;Caiwen Ding;Zhe Li;Geng Yuan;Siyu Liao;Xiaolong Ma;Bo Yuan;Xuehai Qian;Jian Tang;Qinru Qiu;X. Lin
Yanzhi Wang;Caiwen Ding;Zhe Li;Geng Yuan;Siyu Liao;Xiaolong Ma;Bo Yuan;Xuehai Qian;Jian Tang;Qinru Qiu;X. Lin
中科院分区:
其他
文献类型:
--
作者:
Yanzhi Wang;Caiwen Ding;Zhe Li;Geng Yuan;Siyu Liao;Xiaolong Ma;Bo Yuan;Xuehai Qian;Jian Tang;Qinru Qiu;X. Lin

文献摘要

相似文献

深度学习系统的硬件加速已经在工业界和学术界得到了广泛的研究。本文旨在为深度神经网络(DNN)的硬件实现实现超高的能量效率和性能。提出了一种适用于不同DNN类型、规模和应用场景的算法-硬件协同优化框架。算法部分采用了通用的块循环矩阵,实现了精度和压缩比的细粒度折衷。它既适用于完全连通的层,也适用于卷积层,并包含了对该方法有效性的数学严格证明。该算法将训练和推理的每层计算复杂度从O(N2)降低到O(Nlogn),存储复杂度从O(N2)降低到O(N)。硬件部分包括基于现场可编程门阵列(现场可编程门阵列)的高效实现,使用有效的重新配置、批处理、深度流水线、资源重用和分层控制。实验结果表明,与IBM TrueNorth处理器相比,在相同的测试精度下,该框架至少获得了152倍的加速比和71倍的能效提升。与参考的基于现场可编程门阵列的工作相比,它实现了至少31倍的能效提升。
Hardware accelerations of deep learning systems have been extensively investigated in industry and academia. The aim of this paper is to achieve ultra-high energy efficiency and performance for hardware implementations of deep neural networks (DNNs). An algorithm-hardware co-optimization framework is developed, which is applicable to different DNN types, sizes, and application scenarios. The algorithm part adopts the general block-circulant matrices to achieve a fine-grained tradeoff of accuracy and compression ratio. It applies to both fully-connected and convolutional layers and contains a mathematically rigorous proof of the effectiveness of the method. The proposed algorithm reduces computational complexity per layer from O(n2) to O(n log n) and storage complexity from O(n2) to O(n), both for training and inference. The hardware part consists of highly efficient Field Programmable Gate Array (FPGA)-based implementations using effective reconfiguration, batch processing, deep pipelining, resource re-using, and hierarchical control. Experimental results demonstrate that the proposed framework achieves at least 152X speedup and 71X energy efficiency gain compared with IBM TrueNorth processor under the same test accuracy. It achieves at least 31X energy efficiency gain compared with the reference FPGA-based work.