SMART: A Heterogeneous Scratchpad Memory Architecture for Superconductor SFQ-based Systolic CNN Accelerators

SMART: A Heterogeneous Scratchpad Memory Architecture for Superconductor SFQ-based Systolic CNN Accelerators
复制标题

DOI:
10.1145/3466752.3480041
复制
发表时间:
2021-09
期刊:
MICRO-54: 54th Annual IEEE/ACM International Symposium on Microarchitecture
影响因子:
--
通讯作者:
Farzaneh Zokaee;Lei Jiang
Farzaneh Zokaee;Lei Jiang
中科院分区:
其他
文献类型:
--
作者:
Farzaneh Zokaee;Lei Jiang

文献摘要

被引文献

相似文献

基于超快、低功耗超导单通量量子 (SFQ) 的 CNN 脉动加速器旨在增强 CNN 推理吞吐量。然而,由于缺乏随机访问功能,基于移位寄存器 (SHIFT) 的暂存存储器 (SPM) 阵列会阻止 SFQ CNN 加速器超过其峰值吞吐量的 40%。本文首先记录了我们对多种低温存储器技术的研究,包括涡旋过渡存储器(VTM)、约瑟夫森-CMOS SRAM、MRAM和超导纳米线存储器,期间我们发现上述技术都没有使SFQ CNN加速器同时实现高吞吐量、小面积和低功耗。其次,我们提出了一种异构 SPM 架构 SMART,由 SHIFT 阵列和随机存取阵列组成,以提高 SFQ CNN 脉动加速器的推理吞吐量。第三,我们通过构建连接 CMOS 子组的基于 SFQ 无源传输线的 H 树,提出了一种快速、低功耗和密集流水线随机访问 CMOS-SFQ 阵列。最后,我们创建一个基于 ILP 的编译器来在 SMART 上部署 CNN 模型。实验结果表明,在相同的芯片面积开销下,相比最新的基于SHIFT的SFQ CNN加速器,SMART在推断单张图像(一批图像)时,推理吞吐量提高了3.9×(2.2×),推理能量降低了86%(71%)。
Ultra-fast & low-power superconductor single-flux-quantum (SFQ)-based CNN systolic accelerators are built to enhance the CNN inference throughput. However, shift-register (SHIFT)-based scratchpad memory (SPM) arrays prevent a SFQ CNN accelerator from exceeding 40% of its peak throughput, due to the lack of random access capability. This paper first documents our study of a variety of cryogenic memory technologies, including Vortex Transition Memory (VTM), Josephson-CMOS SRAM, MRAM, and Superconducting Nanowire Memory, during which we found that none of the aforementioned technologies made a SFQ CNN accelerator achieve high throughput, small area, and low power simultaneously. Second, we present a heterogeneous SPM architecture, SMART, composed of SHIFT arrays and a random access array to improve the inference throughput of a SFQ CNN systolic accelerator. Third, we propose a fast, low-power and dense pipelined random access CMOS-SFQ array by building SFQ passive-transmission-line-based H-Trees that connect CMOS sub-banks. Finally, we create an ILP-based compiler to deploy CNN models on SMART. Experimental results show that, with the same chip area overhead, compared to the latest SHIFT-based SFQ CNN accelerator, SMART improves the inference throughput by 3.9 × (2.2 ×), and reduces the inference energy by 86% (71%) when inferring a single image (a batch of images).