Mixed-Signal Computing for Deep Neural Network Inference

Mixed-Signal Computing for Deep Neural Network Inference
复制标题

DOI:
10.1109/tvlsi.2020.3020286
复制
发表时间:
2021-01
影响因子:
2.8
通讯作者:
B. Murmann
B. Murmann
中科院分区:
工程技术2区
文献类型:
--
作者:
B. Murmann

文献摘要

被引文献

相似文献

现代深度神经网络(DNN)每个推理需要数十亿次乘法累加运算。考虑到这些计算需要相对较低的精度,考虑模拟计算是可行的,在低SNR状态下,模拟计算比数字计算更有效。这篇概述文章研究了在现代DNN处理器架构的背景下混合模拟/数字计算方法的潜力,这些架构通常受到内存访问的限制。我们将讨论如何内存和内存中的计算结构可能有助于缓解这一瓶颈,并在处理阵列级别获得渐近效率限制。结果表明,对于4位混合信号运算,单位fJ/op能量效率是可行的。在该分析中,特别考虑了模数接口的SNR和摊销要求。此外,我们考虑了各种实现风格的利弊,并强调了为完整的DNN加速器设计保持高计算效率的挑战。
Modern deep neural networks (DNNs) require billions of multiply-accumulate operations per inference. Given that these computations demand relatively low precision, it is feasible to consider analog computing, which can be more efficient than digital in the low-SNR regime. This overview article investigates the potential of mixed analog/digital computing approaches in the context of modern DNN processor architectures, which are typically limited by memory access. We discuss how memory-like and in-memory compute fabrics may help alleviate this bottleneck and derive asymptotic efficiency limits at the processing array level. It is shown that single-digit fJ/op energy efficiencies are feasible for 4-bit mixed-signal arithmetic. In this analysis, special consideration is given to the SNR and amortization requirements of the analog–digital interfaces. In addition, we consider the pros and cons for a variety of implementation styles and highlight the challenge of retaining high compute efficiency for a complete DNN accelerator design.