QUEST: A 7.49TOPS multi-purpose log-quantized DNN inference engine stacked on 96MB 3D SRAM using inductive-coupling technology in 40nm CMOS
QUEST: A 7.49TOPS multi-purpose log-quantized DNN inference engine stacked on 96MB 3D SRAM using inductive-coupling technology in 40nm CMOS
复制标题
DOI:
10.1109/isscc.2018.8310261
复制
发表时间:
2018-03
期刊:
影响因子:
--
通讯作者:
Kodai Ueyoshi;Kota Ando;Kazutoshi Hirose;Shinya Takamaeda-Yamazaki;J. Kadomoto;T. Miyata;M. Hamada;T. Kuroda;M. Motomura
中科院分区:
文献类型:
--
作者:
Kodai Ueyoshi;Kota Ando;Kazutoshi Hirose;Shinya Takamaeda-Yamazaki;J. Kadomoto;T. Miyata;M. Hamada;T. Kuroda;M. Motomura
A key consideration for deep neural network (DNN) inference accelerators is the need for large and high-bandwidth external memories. Although an architectural concept for stacking a DNN accelerator with DRAMs has been proposed previously, long DRAM latency remains problematic and limits the performance [1]. Recent algorithm-level optimizations, such as network pruning and compression, have shown success in reducing the DNN memory size [2]; however, since networks become irregular and sparse, they induce an additional need for agile random accesses to the memory systems.