QUEST: A 7.49TOPS multi-purpose log-quantized DNN inference engine stacked on 96MB 3D SRAM using inductive-coupling technology in 40nm CMOS

QUEST: A 7.49TOPS multi-purpose log-quantized DNN inference engine stacked on 96MB 3D SRAM using inductive-coupling technology in 40nm CMOS
复制标题

DOI:
10.1109/isscc.2018.8310261
复制
发表时间:
2018-03
期刊:
2018 IEEE International Solid - State Circuits Conference - (ISSCC)
影响因子:
--
通讯作者:
Kodai Ueyoshi;Kota Ando;Kazutoshi Hirose;Shinya Takamaeda-Yamazaki;J. Kadomoto;T. Miyata;M. Hamada;T. Kuroda;M. Motomura
Kodai Ueyoshi;Kota Ando;Kazutoshi Hirose;Shinya Takamaeda-Yamazaki;J. Kadomoto;T. Miyata;M. Hamada;T. Kuroda;M. Motomura
中科院分区:
其他
文献类型:
--
作者:
Kodai Ueyoshi;Kota Ando;Kazutoshi Hirose;Shinya Takamaeda-Yamazaki;J. Kadomoto;T. Miyata;M. Hamada;T. Kuroda;M. Motomura

文献摘要

被引文献

相似文献

深度神经网络 (DNN) 推理加速器的一个关键考虑因素是需要大容量、高带宽的外部存储器。尽管之前已经提出了将 DNN 加速器与 DRAM 堆叠的架构概念,但 DRAM 延迟较长仍然存在问题并限制了性能 [1]。最近的算法级优化,例如网络修剪和压缩,已在减少 DNN 内存大小方面取得了成功 [2];然而,由于网络变得不规则和稀疏,它们引发了对存储系统灵活随机访问的额外需求。
A key consideration for deep neural network (DNN) inference accelerators is the need for large and high-bandwidth external memories. Although an architectural concept for stacking a DNN accelerator with DRAMs has been proposed previously, long DRAM latency remains problematic and limits the performance [1]. Recent algorithm-level optimizations, such as network pruning and compression, have shown success in reducing the DNN memory size [2]; however, since networks become irregular and sparse, they induce an additional need for agile random accesses to the memory systems.