Processing-in-Memory Using Optically-Addressed Phase Change Memory

Processing-in-Memory Using Optically-Addressed Phase Change Memory
复制标题

DOI:
10.1109/islped58423.2023.10244409
复制
发表时间:
2023-08
期刊:
2023 IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED)
影响因子:
--
通讯作者:
Guowei Yang;Cansu Demirkıran;Zeynep Ece Kizilates;Carlos A. Ríos Ocampo;Ayse K. Coskun;Ajay Joshi
Guowei Yang;Cansu Demirkıran;Zeynep Ece Kizilates;Carlos A. Ríos Ocampo;Ayse K. Coskun;Ajay Joshi
中科院分区:
其他
文献类型:
--
作者:
Guowei Yang;Cansu Demirkıran;Zeynep Ece Kizilates;Carlos A. Ríos Ocampo;Ayse K. Coskun;Ajay Joshi

文献摘要

相似文献

当今的深度神经网络(DNN)推理系统包含数千亿个参数,由于计算和存储单元之间的频繁数据传输,导致推理过程中的显著延迟和能量开销。内存处理(PiM)已经成为解决这个问题的可行解决方案,避免了昂贵的数据移动。基于电气设备的PiM方法遭受吞吐量和能量效率问题。相比之下,光寻址相变存储器(OPCM)利用光进行操作,并且与其电对应物相比实现了更高的吞吐量和能量效率。本文介绍了一种系统级设计,考虑到OPCM编程开销,并确定编程成本占主导地位的DNN推理基于OPCM的PiM架构。我们探索这个系统的设计空间,并确定最节能的OPCM阵列大小和批量大小。我们提出了一种新的阈值和重排序技术的权重块,以进一步减少编程开销。结合这些优化,我们的方法在实际DNN工作负载中实现了比现有光子加速器高出65.2倍的吞吐量。
Today's Deep Neural Network (DNN) inference systems contain hundreds of billions of parameters, resulting in significant latency and energy overheads during inference due to frequent data transfers between compute and memory units. Processing-in-Memory (PiM) has emerged as a viable solution to tackle this problem by avoiding the expensive data movement. PiM approaches based on electrical devices suffer from throughput and energy efficiency issues. In contrast, Optically-addressed Phase Change Memory (OPCM) operates with light and achieves much higher throughput and energy efficiency compared to its electrical counterparts. This paper introduces a system-level design that takes the OPCM programming overhead into consideration, and identifies that the programming cost dominates the DNN inference on OPCM-based PiM architectures. We explore the design space of this system and identify the most energy-efficient OPCM array size and batch size. We propose a novel thresholding and reordering technique on the weight blocks to further reduce the programming overhead. Combining these optimizations, our approach achieves up to 65.2 × higher throughput than existing photonic accelerators for practical DNN workloads.