Design-Technology Co-Optimization for NVM-Based Neuromorphic Processing Elements

Design-Technology Co-Optimization for NVM-Based Neuromorphic Processing Elements
复制标题

DOI:
10.1145/3524068
复制
发表时间:
2022-03
影响因子:
2
通讯作者:
Shihao Song;Adarsha Balaji;Anup Das;Nagarajan Kandasamy
Shihao Song;Adarsha Balaji;Anup Das;Nagarajan Kandasamy
中科院分区:
计算机科学3区
文献类型:
--
作者:
Shihao Song;Adarsha Balaji;Anup Das;Nagarajan Kandasamy

文献摘要

相似文献

机器学习的新兴用例(ML)是在高性能系统上训练模型,并在能量受限的嵌入式系统上部署训练有素的模型。根据生物大脑的原理运行的神经形态硬件平台可以大大降低ML推理任务的能量开销,从而使这些平台成为嵌入式ML系统的有吸引力的解决方案。我们提出了设计技术权衡分析,以实现基于非挥发记忆(NVM)的神经形态硬件的处理元素(PES)(PES)。通过详细的电路级模拟在缩放过程技术节点上,我们显示了技术缩放对信息处理延迟的负面影响,这会影响嵌入式ML系统的服务质量。在较细的粒度下,PE内的延迟取决于(1)寄生成分在其当前路径上引入的延迟,以及(2)延迟的变化延迟,以感知其NVM细胞的不同电阻状态。基于这两个观察结果,我们做出以下三项贡献。首先,在技术方面,我们提出了一个优化方案,其中NVM阻力状态花费最长的感觉是在当前路径上设置的,反之亦然,从而减少了PE延迟,从而提高了服务质量。其次,在体系结构方面,我们将每个PE中的隔离晶体管介绍为可以单独电源门控的区域,从而减少延迟和能量。最后,在系统软件方面,我们提出了一种机制,可以在实施硬件神经形态PE的ML推理任务时利用所提出的技术和架构增强。通过最近的神经形态硬件结构进行的评估表明,我们提出的设计技术合作方法可以提高ML推理任务的性能和能源效率,而不会导致每位高成本高。
An emerging use case of machine learning (ML) is to train a model on a high-performance system and deploy the trained model on energy-constrained embedded systems. Neuromorphic hardware platforms, which operate on principles of the biological brain, can significantly lower the energy overhead of an ML inference task, making these platforms an attractive solution for embedded ML systems. We present a design-technology tradeoff analysis to implement such inference tasks on the processing elements (PEs) of a non-volatile memory (NVM)-based neuromorphic hardware. Through detailed circuit-level simulations at scaled process technology nodes, we show the negative impact of technology scaling on the information-processing latency, which impacts the quality of service of an embedded ML system. At a finer granularity, the latency inside a PE depends on (1) the delay introduced by parasitic components on its current paths, and (2) the varying delay to sense different resistance states of its NVM cells. Based on these two observations, we make the following three contributions. First, on the technology front, we propose an optimization scheme where the NVM resistance state that takes the longest time to sense is set on current paths having the least delay, and vice versa, reducing the average PE latency, which improves the quality of service. Second, on the architecture front, we introduce isolation transistors within each PE to partition it into regions that can be individually power-gated, reducing both latency and energy. Finally, on the system-software front, we propose a mechanism to leverage the proposed technological and architectural enhancements when implementing an ML inference task on neuromorphic PEs of the hardware. Evaluations with a recent neuromorphic hardware architecture show that our proposed design-technology co-optimization approach improves both performance and energy efficiency of ML inference tasks without incurring high cost-per-bit.