FlutPIM:: A Look-up Table-based Processing in Memory Architecture with Floating-point Computation Support for Deep Learning Applications

FlutPIM:: A Look-up Table-based Processing in Memory Architecture with Floating-point Computation Support for Deep Learning Applications
复制标题

FlutPIM:: 内存架构中基于查找表的处理,支持深度学习应用的浮点计算

DOI:
10.1145/3583781.3590313
复制
发表时间:
2023
期刊:
Great Lakes Symposium on VLSI
影响因子:
--
通讯作者:
Ganguly, Amlan
Ganguly, Amlan
中科院分区:
--
文献类型:
--
作者:
Sutradhar, Purab Ranjan;Bavikadi, Sathwika;Indovina, Mark;Pudukotai Dinakarrao, Sai Manoj;Ganguly, Amlan

文献摘要

参考文献

被引文献

相似文献

内存处理(PIM)在广泛的数据驱动应用中显示出巨大的潜力,特别是深度学习和人工智能。然而,在存储器芯片的有限范围内促进标准处理器(即,CPU或GPU)的计算复杂性而不贡献显著的电路开销是一个挑战。为了解决这一挑战,我们提出了一个可编程的LUT为基础的面积高效的PIM架构,能够执行各种低精度浮点(FP)计算使用一种新的面向LUT的操作数分解技术。我们将这种紧凑的计算单元大量地集成在内存库中,以实现令人印象深刻的并行处理能力,比最先进的具有FP功能的PIM高出4倍。此外,我们采用高度优化的低精度FP格式,可以以最小的计算精度损失最大限度地提高计算性能,尤其是对于深度学习应用程序。与最先进的内存加速技术相比,总体结果是吞吐量提高了17%,每库计算带宽提高了8- 20倍。
Processing-in-Memory (PIM) has shown great potential for a wide range of data-driven applications, especially Deep Learning and AI. However, it is a challenge to facilitate the computational sophistication of a standard processor (i.e. CPU or GPU) within the limited scope of a memory chip without contributing significant circuit overheads. To address the challenge, we propose a programmable LUT-based area-efficient PIM architecture capable of performing various low-precision floating point (FP) computations using a novel LUT-oriented operand-decomposition technique. We incorporate such compact computational units within the memory banks in a large count to achieve impressive parallel processing capabilities, up to 4x higher than state-of-the-art FP-capable PIM. Additionally, we adopt a highly-optimized low-precision FP format that maximizes computational performance at a minimal compromise of computational precision, especially for Deep Learning Applications. The overall result is a 17% higher throughput and an impressive 8-20x higher compute Bandwidth/bank compared to the state-of-the-art of in-memory acceleration.
基于 DRAM 的 PIM 架构优化,用于节能深度神经网络训练
DOI: 10.1109/iscas48785.2022.9937832
发表时间: 2022
期刊: 2022 IEEE International Symposium on Circuits and Systems (ISCAS)
影响因子: --
作者:
C. Sudarshan;Mohammad Hassani Sadi;C. Weis;N. Wehn
通讯作者: N. Wehn
DOI: 10.1109/lca.2020.3011643
发表时间: 2020-07-01
影响因子: 2.3
作者:
Sutradhar, Purab Ranjan;Connolly, Mark;Ganguly, Amlan
通讯作者: Ganguly, Amlan
25.4 20nm 6GB 内存功能 DRAM,基于 HBM2,具有使用组级并行性的 1.2TFLOPS 可编程计算单元,适用于机器学习应用
DOI: --
发表时间: 2021
期刊: IEEE International Solid-State Circuits Conference
影响因子: --
作者:
Young;Suk Han Lee;Jaehoon Lee;Sanghyuk Kwon;Je;J. Son;O. Seongil;Hak;Hae;Sooyoung Kim;Young;Jin Guk Kim;Jo;Hyunsung Shin;J. Kim;BengSeng Phuah;H. Kim;Myeongsoo Song;A. Choi;Daeho Kim;Sooyoung Kim;Eunhwan Kim;David Wang;Shin;Yuhwan Ro;Seungwoo Seo;Joonho Song;Jaeyoun Youn;Kyomin Sohn;N. Kim
通讯作者: N. Kim
缺少内存墙:处理器/内存集成案例
DOI: --
发表时间: 1996
期刊: International Symposium on Computer Architecture
影响因子: --
作者:
Ashley Saulsbury;Fong Pong;A. Nowatzyk
通讯作者: A. Nowatzyk