uSystolic: Byte-Crawling Unary Systolic Array

uSystolic: Byte-Crawling Unary Systolic Array
复制标题

DOI:
10.1109/hpca53966.2022.00010
复制
发表时间:
2022-04
期刊:
2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Di Wu;Joshua San Miguel
Di Wu;Joshua San Miguel
中科院分区:
其他
文献类型:
--
作者:
Di Wu;Joshua San Miguel

文献摘要

相似文献

广义矩阵乘法(GEMM)是广泛应用中的一个重要运算,特别是在蓬勃发展的深度神经网络中。为了实现GEMM的低功耗,研究人员已经利用了一元计算,它用非常简单的逻辑来操纵比特流。然而,现有的一元架构不能很好地推广到通用应用中的不同GEMM配置,并且与二进制计算堆栈不兼容,从而给轻松执行一元GEMM带来了挑战。在这项工作中,我们通过架构一个混合一元-二进制脉动阵列uSyrup来解决这个问题,uSyrup继承了传统的二进制数据调度,数据移动缓慢(因此具有功率效率),即,数据字节从内存中爬出到驱动器uSysControl中。uSYS表现出巨大的面积和功率改进,这是1)低功率计算内核、2)时空比特流重用和3)片上SRAM消除的联合效果。对于评估的边缘计算场景,与二进制并行设计相比,速率编码的uSyrup将脉动阵列面积和总片上面积减少了59.0%和91.3%,对于AlexNet,片上能量和功率效率提高了112.2倍和44.8倍。
General matrix multiply (GEMM) is an important operation in broad applications, especially the thriving deep neural networks. To achieve low power consumption for GEMM, researchers have already leveraged unary computing, which manipulates bitstreams with extremely simple logic. However, existing unary architectures are not well generalizable to varying GEMM configurations in versatile applications and incompatible to the binary computing stack, imposing challenges to execute unary GEMM effortlessly. In this work, we address the problem by architecting a hybrid unary-binary systolic array, uSystolic, to inherit the legacy-binary data scheduling with slow (thus power-efficient) data movement, i.e., data bytes are crawling out from memory to drive uSystolic. uSystolic exhibits tremendous area and power improvements as a joint effect of 1) low-power computing kernel, 2) spatial-temporal bitstream reuse, and 3) on-chip SRAM elimination. For the evaluated edge computing scenario, compared with the binary parallel design, the rated-coded uSystolic reduces the systolic array area and total on-chip area by 59.0% and 91.3%, with the on-chip energy and power efficiency improved by up to 112.2× and 44.8× for AlexNet.