24.2 A 2.5GHz 7.7TOPS/W switched-capacitor matrix multiplier with co-designed local memory in 40nm

24.2 A 2.5GHz 7.7TOPS/W switched-capacitor matrix multiplier with co-designed local memory in 40nm
复制标题

24.2 2.5GHz 7.7TOPS/W 开关电容器矩阵乘法器,具有共同设计的本地存储器,采用 40nm 工艺

DOI:
--
复制
发表时间:
2016
期刊:
IEEE International Solid-State Circuits Conference
影响因子:
--
通讯作者:
S. Wong
S. Wong
中科院分区:
--
文献类型:
--
作者:
Edward H. Lee;S. Wong

文献摘要

被引文献

相似文献

由乘加硬件实现的矩阵乘法在信号处理、计算机图形学、机器学习和优化中无处不在。许多重要的应用具有固有的鲁棒性,以降低矩阵乘法的精度,例如神经网络的推理[1],可以利用模拟信号处理来提高能效。本文提出了一种64周期可编程无源开关电容矩阵乘法器(SCMM),并采用了共同设计的无位线存储器。该设计利用300 aF单位边缘电容进行高速和低能耗电荷域处理,并包含输入DAC、乘法累加SAR ADC和本地存储器。SCMM的两个应用被证明:1)图像分类器系统的模拟前端,它减少了21倍的A/D转换和11倍的乘法和累加计算能量比传统的系统,和2)协处理加速器解决随机梯度下降优化,它实现了测量7.7 TOPS/W在2.5GHz。
Matrix multiplication, enabled by multiply-and-accumulate hardware, is ubiquitous in signal processing, computer graphics, machine learning, and optimization. Many important applications with inherent robustness to reduced precision for matrix multiplication, e.g. inference for neural networks [1], can take advantage of analog signal processing for energy efficiency. This work presents a 64-cycle programmable passive Switched-Capacitor Matrix Multiplier (SCMM) with co-designed bitline-less memory. The design exploits 300aF unit fringe capacitors for high speed and low energy charge-domain processing and contains the input DAC, multiply-and-accumulate SAR ADC, and local memory. Two applications of the SCMM are demonstrated: 1) an analog front-end for an image classifier system, which reduces A/D conversions by 21x and multiply-and-accumulate compute energy by 11x over a conventional system, and 2) a co-processing accelerator to solve Stochastic Gradient Descent optimization, which achieves a measured 7.7 TOPS/W at 2.5GHz.