Benchmark of RRAM based Architectures for Dot-Product Computation

Benchmark of RRAM based Architectures for Dot-Product Computation
复制标题

DOI:
10.1109/apccas.2018.8605606
复制
发表时间:
2018-10
期刊:
2018 IEEE Asia Pacific Conference on Circuits and Systems (APCCAS)
影响因子:
--
通讯作者:
Xiaochen Peng;Shimeng Yu
Xiaochen Peng;Shimeng Yu
中科院分区:
其他
文献类型:
--
作者:
Xiaochen Peng;Shimeng Yu

文献摘要

相似文献

基于新兴非易失性存储器件的存储阵列架构已被提出,用于神经网络中内积计算的片上加速。由于机器学习的最新进展表明,精度降低是一种减少计算和存储的有用技术,因此有必要评估其硬件成本。在本文中,我们使用一种电路级宏模型,即NeuroSim,对XNOR - RRAM和传统8位RRAM架构的电路级性能指标进行基准测试,这些指标包括芯片面积、延迟和动态能量。这两种架构分别以逐行顺序和并行读出的方式实现对512×512突触矩阵的内积运算。模拟结果基于RRAM模型和32nm CMOS工艺设计套件(PDK),并行XNOR - RRAM架构的能效可达到311万亿次运算每秒每瓦(TOPS/W),与并行和顺序的传统8位RRAM架构相比,分别显示出至少约15倍和约621倍的提升。
Memory array architecture based on emerging non-volatile memory devices have been proposed for on-chip acceleration of dot-product computation in neural networks. As recent advances in machine learning have shown that precision reduction is a useful technique to reduce the computation and memory storage, it is desired to evaluate their hardware cost. In this paper, we use a circuit-level macro model, i.e. NeuroSim, to benchmark the circuit-level performance metrics, such as chip area, latency, and dynamic energy for the XNOR-RRAM and conventional 8-bit RRAM architectures. Both architectures are implemented to process the dot-product operation of a 512×512 synaptic matrix in sequential row-by-row and parallel read-out fashion separately. The simulation results are based on RRAM models and 32nm CMOS PDK, the energy-efficiency of the parallel XNOR-RRAM architecture could achieve 311 TOPS/W, showing at least ~15× and ~621× improvement compared to the parallel and sequential conventional 8-bit RRAM architectures respectively.