25.4 A 20nm 6GB Function-In-Memory DRAM, Based on HBM2 with a 1.2TFLOPS Programmable Computing Unit Using Bank-Level Parallelism, for Machine Learning Applications

25.4 A 20nm 6GB Function-In-Memory DRAM, Based on HBM2 with a 1.2TFLOPS Programmable Computing Unit Using Bank-Level Parallelism, for Machine Learning Applications
复制标题

25.4 20nm 6GB 内存功能 DRAM,基于 HBM2,具有使用组级并行性的 1.2TFLOPS 可编程计算单元,适用于机器学习应用

DOI:
--
复制
发表时间:
2021
期刊:
IEEE International Solid-State Circuits Conference
影响因子:
--
通讯作者:
N. Kim
N. Kim
中科院分区:
--
文献类型:
--
作者:
Young;Suk Han Lee;Jaehoon Lee;Sanghyuk Kwon;Je;J. Son;O. Seongil;Hak;Hae;Sooyoung Kim;Young;Jin Guk Kim;Jo;Hyunsung Shin;J. Kim;BengSeng Phuah;H. Kim;Myeongsoo Song;A. Choi;Daeho Kim;Sooyoung Kim;Eunhwan Kim;David Wang;Shin;Yuhwan Ro;Seungwoo Seo;Joonho Song;Jaeyoun Youn;Kyomin Sohn;N. Kim

文献摘要

被引文献

相似文献

近年来,人工智能(AI)技术迅速扩散,广泛应用于语音识别、医疗保健和自动驾驶等应用领域。为了提高人工智能的能力,需要更强大的系统来处理更大量的数据。这一要求使得特定领域的加速器(如GPU和TPU)很受欢迎,因为它们可以提供比最先进的CPU高几个数量级的性能。然而,这些加速器只有在从存储器中获取必要的数据时才能以最高性能运行:需要具有高带宽和大容量的片外存储器[1]。到目前为止,HBM已经满足了带宽和容量要求[2] -[6],但最近的人工智能技术,如递归神经网络,需要比HBM更高的带宽[7]-[8]。虽然片外带宽的进一步增加可以通过各种技术来实现,但它通常受到芯片或系统级功率约束的限制[9]。因此,降低非常规架构(如内存处理)对片外带宽的需求至关重要。在本文中,我们提出了功能在内存DRAM(FIMDRAM),集成了一个16宽的单指令多数据引擎的内存银行和利用银行级并行提供4美元 处理带宽比片外存储器解决方案高10倍。其次,我们展示了不需要对传统存储器控制器及其命令协议进行任何修改的技术,这使得FIMDRAM更实用,可快速在行业中采用。最后,我们得出结论,本文与电路和系统级的评价,我们制造的FIMDRAM。
In recent years, artificial intelligence (AI) technology has proliferated rapidly and widely into application areas such as speech recognition, health care, and autonomous driving. To increase the capabilities of AI more powerful systems are needed to process a larger amount of data. This requirement has made domain-specific accelerators, such as GPUs and TPUs, popular; as they can provide orders of magnitude higher performance than state-of-the-art CPUs. However, these accelerators can only operate at their peak performance when they get the necessary data from memory as quickly as it is processed: requiring off-chip memory with a high bandwidth and a large capacity [1]. HBM has thus far met the bandwidth and capacity requirement [2] –[6], but recent AI technologies such as recurrent neural networks require an even higher bandwidth than HBM [7]–[8]. While a further increase in off-chip bandwidth can be accomplished by various techniques, it is often limited by power constraints at the chip or system level [9]. Hence, it is essential to decrease demand for off-chip bandwidth with unconventional architectures: such as processing-in-memory. In this paper, we present function-In-memory DRAM (FIMDRAM) that integrates a 16-wide single-instruction multiple-data engine within the memory banks and that exploits bank-level parallelism to provide $4 imes $ higher processing bandwidth than an off-chip memory solution. Second, we show techniques that do not require any modification to conventional memory controllers and their command protocols, which make FIMDRAM more practical for quick industry adoption. Finally, we conclude this paper with circuit- and system-level evaluations of our fabricated FIMDRAM.