pPIM: A Programmable Processor-in-Memory Architecture With Precision-Scaling for Deep Learning

pPIM: A Programmable Processor-in-Memory Architecture With Precision-Scaling for Deep Learning
复制标题

DOI:
10.1109/lca.2020.3011643
复制
发表时间:
2020-07-01
影响因子:
2.3
通讯作者:
Ganguly, Amlan
Ganguly, Amlan
中科院分区:
计算机科学3区
文献类型:
--
作者:
Sutradhar, Purab Ranjan;Connolly, Mark;Ganguly, Amlan

文献摘要

被引文献

相似文献

存储器访问延迟和低数据传输带宽限制了许多数据密集型应用的处理速度,例如传统冯·诺伊曼体系结构中的卷积神经网络(CNN)。内存中处理(PIM)被认为是一种潜在的硬件解决方案,因为PIM中的数据访问瓶颈可以通过在内存芯片内执行计算来避免。然而,在存储器内使用基于逻辑的复杂处理单元来实现PIM带来了复杂的制造挑战。在这封信中,我们建议利用现有的存储器基础设施来实现可编程PIM(PPIM),这是一种基于查找表(LUT)的新型PIM,其中所有处理单元都只使用LUT实现,而不是以前基于LUT的PIM实现将LUT与逻辑电路相结合进行计算。这使PPIM能够以最小的制造复杂性执行超低功耗和低延迟操作。此外,完整的基于LUT的设计在PPIM中提供了简单的基于存储器写入的可编程性。启用精确缩放功能可进一步提高CNN应用的性能和功耗。可编程性功能可能会使在线培训的实施变得更容易。我们的初步模拟表明,与现有的传统处理器结构、图形处理单元(GPU)和基于混合查找表逻辑的PIM相比,我们提出的PPIM在推理吞吐量上分别提高了2000倍、657.5倍和1.46倍。此外,精确缩放将PPIM的能效提高了约1.35倍,超过其全精度操作。
Memory access latencies and low data transfer bandwidth limit the processing speed of many data intensive applications such as Convolutional Neural Networks (CNNs) in conventional Von Neumann architectures. Processing in Memory (PIM) is envisioned as a potential hardware solution for such applications as the data access bottlenecks can be avoided in PIM by performing computations within the memory die. However, PIM realizations with logic-based complex processing units within the memory present complicated fabrication challenges. In this letter, we propose to leverage the existing memory infrastructure to implement a programmable PIM (pPIM), a novel Look-Up-Table (LUT)-based PIM where all the processing units are implemented solely with LUTs, as opposed to prior LUT-based PIM implementations that combine LUT with logic circuitry for computations. This enables pPIM to perform ultra-low power & low-latency operations with minimal fabrication complications. Moreover, the complete LUT-based design offers simple 'memory write' based programmability in pPIM. Enabling precision scaling further improves the performance and the power consumption for CNN applications. The programmability feature potentially makes it easier for online training implementations. Our preliminary simulations demonstrate that our proposed pPIM can achieve 2000x, 657.5x and 1.46x improvement in inference throughput per unit power consumption compared to state-of-the-art conventional processor architecture, Graphics Processing Unit (GPUs) and a prior hybrid LUT-logic based PIM respectively. Furthermore, precision scaling improves the energy efficiency of the pPIM approximately by 1.35x over its full-precision operation.