A Stochastic-Computing based Deep Learning Framework using Adiabatic Quantum-Flux-Parametron Superconducting Technology

A Stochastic-Computing based Deep Learning Framework using Adiabatic Quantum-Flux-Parametron Superconducting Technology
复制标题

DOI:
10.1145/3307650.3322270
复制
发表时间:
2019-06
期刊:
2019 ACM/IEEE 46th Annual International Symposium on Computer Architecture (ISCA)
影响因子:
--
通讯作者:
R. Cai;Ao Ren;O. Chen;Ning Liu;Caiwen Ding;Xuehai Qian;Jie Han;Wenhui Luo;N. Yoshikawa;Yanzhi Wang
R. Cai;Ao Ren;O. Chen;Ning Liu;Caiwen Ding;Xuehai Qian;Jie Han;Wenhui Luo;N. Yoshikawa;Yanzhi Wang
中科院分区:
其他
文献类型:
--
作者:
R. Cai;Ao Ren;O. Chen;Ning Liu;Caiwen Ding;Xuehai Qian;Jie Han;Wenhui Luo;N. Yoshikawa;Yanzhi Wang

文献摘要

被引文献

相似文献

绝热量子磁通参变(AQFP)超导技术是近年来发展起来的一种新型超导技术,在超导逻辑家族中具有最高的能效,与目前最先进的CMOS机相比有104-105的潜在增益。2016年,基于AQFP的8.3万JJ规模电路的制造和测试成功,展示了使用AQFP实现大规模系统的可扩展性和潜力。因此,AQFP在高性能计算和深空间应用中具有广阔的应用前景,深度神经网络(DNN)推理加速就是一个重要的例子。除了超高的能效,AQFP还表现出两个独特的特点:深流水线性质,因为每个AQFP逻辑门都连接一个交流时钟信号,这增加了避免原始危险的难度;第二是使用单个AQFP缓冲器产生真正的随机数(RNG)的独特机会,远远高于CMOS中的RNG。我们指出,这两个特性使得AQFP特别兼容随机计算(SC)技术,该技术使用与时间无关的比特序列来表示值,并与深度流水线性质兼容。此外,在前人的工作中已经研究了SC在DNN中的应用,并说明了SC的适用性,因为SC与近似计算更兼容。这是首次使用AQFP技术开发基于SC的DNN加速框架。AQFP电路的深度流水线特性决定了AQFP中累加器/计数器的设计难度,这使得基于SC的DNN中已有的设计不再适用。我们克服了这一限制,考虑了卷积层和FC层的不同性质:(I)FC层的内积计算比Conv层具有更多的输入;(Ii)准确的激活函数在Conv层而不是FC层中至关重要。基于这些观察结果,我们提出了(I)利用双调排序网络和反馈环在卷积层中精确整合求和函数和激活函数,以及(Ii)基于多数门链的低复杂度FC层分类块。为了完成设计,我们还开发了(I)AQFP中的超高效随机数发生器,(Ii)AQFP中的高精度子采样(池化)模块,以及(Iii)为进一步提高性能和自动插入缓冲器/分离器以满足AQFP电路要求的多数综合。实验结果表明,在MNIST数据集上保持96%的准确率的情况下,使用AQFP的基于SC的DNN可以获得高达6.8×104倍的能量效率。
The Adiabatic Quantum-Flux-Parametron (AQFP) superconducting technology has been recently developed, which achieves the highest energy efficiency among superconducting logic families, potentially 104–105 gain compared with state-of-the-art CMOS. In 2016, the successful fabrication and testing of AQFP-based circuits with the scale of 83,000 JJs have demonstrated the scalability and potential of implementing large-scale systems using AQFP. As a result, it will be promising for AQFP in high-performance computing and deep space applications, with Deep Neural Network (DNN) inference acceleration as an important example. Besides ultra-high energy efficiency, AQFP exhibits two unique characteristics: the deep pipelining nature since each AQFP logic gate is connected with an AC clock signal, which increases the difficulty to avoid RAW hazards; the second is the unique opportunity of true random number generation (RNG) using a single AQFP buffer, far more efficient than RNG in CMOS. We point out that these two characteristics make AQFP especially compatible with the stochastic computing (SC) technique, which uses a time-independent bit sequence for value representation, and is compatible with the deep pipelining nature. Further, the application of SC has been investigated in DNNs in prior work, and the suitability has been illustrated as SC is more compatible with approximate computations. This work is the first to develop an SC-based DNN acceleration framework using AQFP technology. The deep-pipelining nature of AQFP circuits translates into the difficulty in designing accumulators/counters in AQFP, which makes the prior design in SC-based DNN not suitable. We overcome this limitation taking into account different properties in CONV and FC layers: (i) the inner product calculation in FC layers has more number of inputs than that in CONV layers; (ii) accurate activation function is critical in CONV rather than FC layers. Based on these observations, we propose (i) accurate integration of summation and activation function in CONV layers using bitonic sorting network and feedback loop, and (ii) low-complexity categorization block for FC layers based on chain of majority gates. For complete design we also develop (i) ultra-efficient stochastic number generator in AQFP, (ii) a high-accuracy sub-sampling (pooling) block in AQFP, and (iii) majority synthesis for further performance improvement and automatic buffer/splitter insertion for requirement of AQFP circuits. Experimental results suggest that the proposed SC-based DNN using AQFP can achieve up to 6.8 × 104 times higher energy efficiency compared to CMOS-based implementation while maintaining 96% accuracy on the MNIST dataset.