A 1.40mm2 141mW 898GOPS sparse neuromorphic processor in 40nm CMOS

A 1.40mm2 141mW 898GOPS sparse neuromorphic processor in 40nm CMOS
复制标题

采用 40nm CMOS 的 1.40mm2 141mW 898GOPS 稀疏神经拟态处理器

DOI:
--
复制
发表时间:
2016
期刊:
2016 IEEE Symposium on VLSI Circuits (VLSI-Circuits)
影响因子:
--
通讯作者:
Zhengya Zhang
Zhengya Zhang
中科院分区:
--
文献类型:
--
作者:
Phil C. Knag;Chester Liu;Zhengya Zhang

文献摘要

被引文献

相似文献

稀疏性是一种受大脑启发的属性,可以显著减少深度学习的工作量和功耗。本文提出了一种1.40mm2 40nm CMOS稀疏神经形态处理器,实现了一个两层卷积限制玻尔兹曼机(CRBM)的推理和支持向量机(SVM)分类器。该处理器采用稀疏卷积器实现稀疏比例的工作负载减少。该架构是并行沿着非稀疏的维度,以尽量减少停滞。在0.9V和240MHz下,该处理器实现了898.2GOPS的有效性能,功耗为140.9mW。使用稀疏性,我们将工作负载、数据路径功耗和面积分别减少了3.4倍、3.3倍和1.74倍。该设计使用基于锁存器的存储器来减少面积,并使用动态时钟门控来节省功耗。
Sparsity is a brain-inspired property that enables a significant reduction in workload and power dissipation of deep learning. This work presents a 1.40mm2 40nm CMOS sparse neuromorphic processor that implements a two-layer convolutional restricted Boltzmann machine (CRBM) for inference and a support vector machine (SVM) classifier. The processor incorporates sparse convolvers to realize sparsity-proportional workload reduction. The architecture is parallelized along a non-sparse dimension to minimize stalling. At 0.9V and 240MHz, the processor achieves an effective 898.2GOPS performance, dissipating 140.9mW. Using sparsity, we reduce the workload, datapath power consumption and area by 3.4×, 3.3× and 1.74×, respectively. The design uses latch-based memory to reduce area and dynamic clock gating to save power.