Inference engine benchmarking across technological platforms from CMOS to RRAM

Inference engine benchmarking across technological platforms from CMOS to RRAM
复制标题

DOI:
10.1145/3357526.3357566
复制
发表时间:
2019-09
期刊:
Proceedings of the International Symposium on Memory Systems
影响因子:
--
通讯作者:
Xiaochen Peng;Minkyu Kim;Xiaoyu Sun;Shihui Yin;Titash Rakshit;R. Hatcher;J. Kittl;Jae-sun Seo;Shimeng Yu
Xiaochen Peng;Minkyu Kim;Xiaoyu Sun;Shihui Yin;Titash Rakshit;R. Hatcher;J. Kittl;Jae-sun Seo;Shimeng Yu
中科院分区:
其他
文献类型:
--
作者:
Xiaochen Peng;Minkyu Kim;Xiaoyu Sun;Shihui Yin;Titash Rakshit;R. Hatcher;J. Kittl;Jae-sun Seo;Shimeng Yu

文献摘要

被引文献

相似文献

最先进的深度卷积神经网络(CNN)被广泛应用于当前的人工智能系统中,并在图像/语音识别和分类方面取得了显着的成功。最近的许多努力已经尝试基于各种方法来设计定制推理引擎,包括脉动架构、近存储器处理和具有诸如电阻式随机存取存储器(RRAM)的新兴技术的存储器中处理(PIM)方法。然而,在统一的框架内对这些不同的方法进行全面比较的情况并不存在,而且新设计或新兴技术的好处大多基于定性预测。在本文中,我们在CIFAR-10数据集上评估了VGG类CNN推理加速器的能效和帧速率,这些数据集跨越了从CMOS到后CMOS的技术平台,具有硬件资源约束,即可比的片上面积。我们还研究了芯片外存储器DRAM访问和互连在数据移动过程中的影响,这是CMOS平台的瓶颈。我们的定量分析表明,外围设备(ADC)占主导地位的能量消耗和面积(而不是存储器阵列)的数字RRAM为基础的并行读出PIM架构。尽管存在ADC,但这种架构的能效(TOPS/W)比脉动阵列或近存储器处理提高了2.5倍以上,由于减少了DRAM访问、高吞吐量和优化的并行读出,帧速率相当。通过减少位数的XNOR网络和流水线技术,可以进一步提高性能,提高10倍以上。
State-of-the-art deep convolutional neural networks (CNNs) are widely used in current AI systems, and achieve remarkable success in image/speech recognition and classification. A number of recent efforts have attempted to design custom inference engine based on various approaches, including the systolic architecture, near memory processing, and processing-in-memory (PIM) approach with emerging technologies such as resistive random access memory (RRAM). However, a comprehensive comparison of these various approaches in a unified framework is missing, and the benefits of new designs or emerging technologies are mostly based on qualitative projections. In this paper, we evaluate the energy efficiency and frame rate for a VGG-like CNN inference accelerator on CIFAR-10 dataset across the technological platforms from CMOS to post-CMOS, with hardware resource constraint, i.e. comparable on-chip area. We also investigate the effects of off-chip memory DRAM access and interconnect during data movement, which are the bottlenecks of CMOS platforms. Our quantitative analysis shows that the peripheries (ADCs) dominate in energy consumption and area (rather than memory array) in digital RRAM-based parallel readout PIM architecture. Despite presence of ADCs, this architecture shows >2.5× improvement in energy efficiency (TOPS/W) over systolic arrays or near memory processing, with a comparable frame rate due to reduced DRAM access, high throughput and optimized parallel read out. Further >10× improvements can be achieved by implementing bit-count reduced XNOR network and pipelining.