Area and Energy Optimization for Bit-Serial Log-Quantized DNN Accelerator with Shared Accumulators

Area and Energy Optimization for Bit-Serial Log-Quantized DNN Accelerator with Shared Accumulators
复制标题

DOI:
10.1109/mcsoc2018.2018.00048
复制
发表时间:
2018-09
期刊:
2018 IEEE 12th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC)
影响因子:
--
通讯作者:
Takumi Kudo;Kodai Ueyoshi;Kota Ando;Kazutoshi Hirose;Ryota Uematsu;Yuka Oba;M. Ikebe;T. Asai;M. Motomura;Shinya Takamaeda-Yamazaki
Takumi Kudo;Kodai Ueyoshi;Kota Ando;Kazutoshi Hirose;Ryota Uematsu;Yuka Oba;M. Ikebe;T. Asai;M. Motomura;Shinya Takamaeda-Yamazaki
中科院分区:
其他
文献类型:
--
作者:
Takumi Kudo;Kodai Ueyoshi;Kota Ando;Kazutoshi Hirose;Ryota Uematsu;Yuka Oba;M. Ikebe;T. Asai;M. Motomura;Shinya Takamaeda-Yamazaki

文献摘要

相似文献

在深度神经网络(DNN)的显著发展中,迫切需要开发一种高度优化的DNN加速器,用于边缘计算,同时具有更少的硬件资源和高计算性能。作为一个众所周知的特性,DNN处理涉及大量的乘法和累加运算。因此,低精度量化,如二进制和对数,是边缘计算设备中具有严格限制的电路资源和能量的必要技术。量化中的位宽要求取决于应用特性。基于位串行处理的可变位宽架构已被提出作为一种可扩展的替代方案,通过统一的硬件结构允许不同的性能和精度平衡要求。在本文中,我们提出了一个优化的DNN硬件架构,支持二进制和可变位宽对数量化。其核心思想是分布式和共享累加器,它通过一个累加器处理多个位串行输入,并为二进制模式提供额外的低开销电路。评估结果表明,与现有架构相比,该想法减少了29.8%的硬件资源,而不会损失任何功能,计算速度和识别精度。此外,使用VGG 16的实用DNN模型,它实现了19.6%的能耗降低。
In the remarkable evolution of deep neural network (DNN), development of a highly optimized DNN accelerator for edge computing with both less hardware resource and high computing performance is strongly required. As a well-known characteristic, DNN processing involves a large number multiplication and accumulation operations. Thus, low-precision quantization, such as binary and logarithm, is an essential technique in edge computing devices with strict restriction of circuit resource and energy. Bit-width requirement in quantization depends on application characteristics. Variable bit-width architecture based on the bit-serial processing has been proposed as a scalable alternative that allows different requirements of performance and accuracy balance by a unified hardware structure. In this paper, we propose a well-optimized DNN hardware architecture with supports of binary and variable bit-width logarithmic quantization. The key idea is the distributed-and-shared accumulator that processes multiple bit-serial inputs by a single accumulator with an additional low-overhead circuit for the binary mode. The evaluation results show that the idea reduces hardware resources by 29.8% compared to the prior architecture without losing any functionality, computing speed, and recognition accuracy. Moreover, it achieves 19.6% energy reduction using a practical DNN model of VGG 16.