Heterogeneous Multi-Functional Look-Up-Table-based Processing-in-Memory Architecture for Deep Learning Acceleration

Heterogeneous Multi-Functional Look-Up-Table-based Processing-in-Memory Architecture for Deep Learning Acceleration
复制标题

DOI:
10.1109/isqed57927.2023.10129338
复制
发表时间:
2023-04
期刊:
2023 24th International Symposium on Quality Electronic Design (ISQED)
影响因子:
--
通讯作者:
Sathwika Bavikadi;Purab Ranjan Sutradhar;A. Ganguly;Sai Manoj Pudukotai Dinakarrao
Sathwika Bavikadi;Purab Ranjan Sutradhar;A. Ganguly;Sai Manoj Pudukotai Dinakarrao
中科院分区:
其他
文献类型:
--
作者:
Sathwika Bavikadi;Purab Ranjan Sutradhar;A. Ganguly;Sai Manoj Pudukotai Dinakarrao

文献摘要

被引文献

相似文献

深度神经网络 (DNN) 和卷积神经网络 (CNN) 等新兴应用使用大量数据来执行计算和数据分析。此类应用程序通常会导致资源限制,并在内存和计算单元之间的数据移动中施加大量开销。为了缓解传统计算架构的带宽瓶颈和低效率,引入了内存处理(PIM)等多种架构。然而,现有的 PIM 架构代表了功耗、性能、面积、能源效率和可编程性之间的权衡。为了更好地在硬件加速器中同时实现能效和灵活性标准,我们在这项工作中引入了基于多功能查找表(LUT)的可重构 PIM 架构。所提出的架构是一种多核架构,每个核心都包含处理元件(PE),这是一个具有使用高速可重配置LUT构建的可编程功能单元的独立处理器。所提出的 LUT 可以执行各种操作,包括 CNN 加速所需的卷积、池化和激活。此外,所提出的LUT能够同时提供与不同功能相关的多个输出,而不需要为不同的功能设计不同的LUT。这可以优化面积和功率开销。此外,我们还设计了特殊功能LUT,它可以提供乘法和累加的同时输出,以及双曲线和S形等特殊激活函数。我们评估了各种 CNN,例如 LeNet、AlexNet 和 ResNet18、34、50。我们的实验结果表明,当在所提出的架构上实现 AlexNet 时,与基于 DRAM 的基于 LUT 的 PIM 架构相比,能源效率最高提高 200 倍,吞吐量提高 1.5 倍。
Emerging applications including deep neural networks (DNNs) and convolutional neural networks (CNNs) employ massive amounts of data to perform computations and data analysis. Such applications often lead to resource constraints and impose large overheads in data movement between memory and compute units. Several architectures such as Processing-in-Memory (PIM) are introduced to alleviate the bandwidth bottlenecks and inefficiency of traditional computing architectures. However, the existing PIM architectures represent a trade-off between power, performance, area, energy efficiency, and programmability. To better achieve the energy-efficiency and flexibility criteria simultaneously in hardware accelerators, we introduce a multi-functional look-up-table (LUT)-based reconfigurable PIM architecture in this work. The proposed architecture is a many-core architecture, each core comprises processing elements (PEs), a stand-alone processor with programmable functional units built using high-speed reconfigurable LUTs. The proposed LUTs can perform various operations, including convolutional, pooling, and activation that are required for CNN acceleration. Additionally, the proposed LUTs are capable of providing multiple outputs relating to different functionalities simultaneously without the need to design different LUTs for different functionalities. This leads to optimized area and power overheads. Furthermore, we also design special-function LUTs, which can provide simultaneous outputs for multiplication and accumulation as well as special activation functions such as hyperbolics and sigmoids. We have evaluated various CNNs such as LeNet, AlexNet, and ResNet18,34,50. Our experimental results have demonstrated that when AlexNet is implemented on the proposed architecture shows a maximum of 200× higher energy efficiency and 1.5× higher throughput than a DRAM-based LUT-based PIM architecture.