Accelerating hybrid and compact neural networks targeting perception and control domains with coarse-grained dataflow reconfiguration

Accelerating hybrid and compact neural networks targeting perception and control domains with coarse-grained dataflow reconfiguration
复制标题

通过粗粒度数据流重新配置加速针对感知和控制领域的混合和紧凑神经网络

DOI:
10.1088/1674-4926/41/2/022401
复制
发表时间:
2020-02
期刊:
Chinese Journal of Semiconductors
影响因子:
--
通讯作者:
Z.Yu
Z.Yu
中科院分区:
其他
文献类型:
--
作者:
Z.Wang;L.Zhou;W.Xie;W.Chen;J.Su;W.Chen;A.Du;S.Li;M.Liang;Y.Lin;W.Zhao;Y.Wu;孙天夫;W.Fang;Z.Yu

文献摘要

参考文献

被引文献

相似文献

在纳米半导体技术不断扩展的推动下,过去几年见证了机器学习技术和应用的逐步发展。最近,专用的机器学习加速器,特别是神经网络,吸引了计算机架构师和VLSI设计师的研究兴趣。最先进的加速器通过部署大量的处理元件来提高性能,但是仍然面临混合和非标准算法内核的资源利用率降低的问题。在这项工作中,我们利用重要的神经网络内核的感知和控制的属性,提出了一个可重构的低功耗处理器,根据网络内核调整数据流的模式,处理元件和片上存储器的功能。与现有的细粒度数据流技术相比,本文提出的粗粒度数据流重构方法能够实现计算和存储资源的广泛共享。构建了三个混合网络,用于MobileNet,深度强化学习和序列分类,并使用定制的指令集和工具链进行分析。采用UMC65nmCMOS工艺设计并制作了一个测试芯片,在1.8 × 1.8mm2的芯片上,在100MHz频率下的功耗为7.51mW。
Driven by continuous scaling of nanoscale semiconductor technologies, the past years have witnessed the progressive advancement of machine learning techniques and applications. Recently, dedicated machine learning accelerators, especially for neural networks, have attracted the research interests of computer architects and VLSI designers. State-of-the-art accelerators increase performance by deploying a huge amount of processing elements, however still face the issue of degraded resource utilization across hybrid and non-standard algorithmic kernels. In this work, we exploit the properties of important neural network kernels for both perception and control to propose a reconfigurable dataflow processor, which adjusts the patterns of data flowing, functionalities of processing elements and on-chip storages according to network kernels. In contrast to state-of-the-art fine-grained data flowing techniques, the proposed coarse-grained dataflow reconfiguration approach enables extensive sharing of computing and storage resources. Three hybrid networks for MobileNet, deep reinforcement learning and sequence classification are constructed and analyzed with customized instruction sets and toolchain. A test chip has been designed and fabricated under UMC 65 nm CMOS technology, with the measured power consumption of 7.51 mW under 100 MHz frequency on a die size of 1.8 × 1.8 mm2.
DOI: 10.1109/apccas.2018.8605639
发表时间: 2018-10
期刊: 2018 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS)
影响因子: --
作者:
Minglan Liang;Mingsong Chen;Zheng Wang;Jingwei Sun
通讯作者: Minglan Liang;Mingsong Chen;Zheng Wang;Jingwei Sun
DOI: 10.1109/tnn.1998.712192
发表时间: 1998
期刊: IEEE Trans. Neural Networks
影响因子: --
作者:
R. S. Sutton;A. Barto
通讯作者: R. S. Sutton;A. Barto
DOI: --
发表时间: 2016-02
期刊: ArXiv
影响因子: --
作者:
F. Iandola;Matthew W. Moskewicz;Khalid Ashraf;Song Han;W. Dally;K. Keutzer
通讯作者: F. Iandola;Matthew W. Moskewicz;Khalid Ashraf;Song Han;W. Dally;K. Keutzer
DOI: 10.1162/089976600300015015
发表时间: 2000-10-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
Gers, FA;Schmidhuber, J;Cummins, F
通讯作者: Cummins, F
DOI: 10.1007/s11263-015-0816-y
发表时间: 2015-12-01
影响因子: 19.5
作者:
Russakovsky, Olga;Deng, Jia;Fei-Fei, Li
通讯作者: Fei-Fei, Li