Pod-racing: bulk-bitwise to floating-point compute in racetrack memory for machine learning at the edge

Pod-racing: bulk-bitwise to floating-point compute in racetrack memory for machine learning at the edge
复制标题

DOI:
10.1109/mm.2022.3195761
复制
发表时间:
2022-11
期刊:
影响因子:
3.6
通讯作者:
S. Ollivier;Xinyi Zhang;Yue Tang;C. Choudhuri;Jingtong Hu;A.K. Jones
S. Ollivier;Xinyi Zhang;Yue Tang;C. Choudhuri;Jingtong Hu;A.K. Jones
中科院分区:
计算机科学3区
文献类型:
--
作者:
S. Ollivier;Xinyi Zhang;Yue Tang;C. Choudhuri;Jingtong Hu;A.K. Jones

文献摘要

相似文献

卷积神经网络(CNN)已成为一种无处不在的算法,在移动和边缘设置中应用程序不断增长横向读取,一种可以确定多个相邻域中“ 1”数量的技术,Pod-Racing可以有效地实现多方面的散装和加法计算和两项乘法。我们证明了使用RM CIM进行反向传播的几个CNN的实施,并将其与CNN推理和培训的最新实现进行了比较培训,pod-racing将效率提高了2倍,能源消耗$ \ geq $≥27%,而$ \ geq $≥18%的吞吐量与最先进的现场可编程栅极阵列加速器相比。
Convolutional neural networks (CNNs) have become a ubiquitous algorithm with growing applications in mobile and edge settings. We describe a compute-in-memory (CIM) technique called POD-RACING using Racetrack memory (RM) to accelerate CNNs for edge systems. Using transverse read, a technique that can determine the number of “1”s in multiple adjacent domains, POD-RACING can efficiently implement multioperand bulk-bitwise and addition computations, and two-operand multiplication. We discuss how POD-RACING can implement both variable precision integer and floating point arithmetic using digital CIM. This allows both CNN inference and on-device training without expensive data movement to the cloud. Based on these functions we demonstrate the implementation of several CNNs with backpropagation using RM CIM and compare these to the state-of-the-art implementations of CNN inference and training. During training, POD-RACING improves efficiency by 2×, energy consumption by $\geq$≥27%, and increases throughput by $\geq$≥18% versus a state-of-the-art field-programmable gate array accelerator.