Design and Analysis of a Nano-photonic Processing Unit for Low-Latency Recurrent Neural Network Applications

Design and Analysis of a Nano-photonic Processing Unit for Low-Latency Recurrent Neural Network Applications
复制标题

DOI:
10.1109/mcsoc57363.2022.00058
复制
发表时间:
2022-12
期刊:
2022 IEEE 15th International Symposium on Embedded Multicore/Many-core Systems-on-Chip (MCSoC)
影响因子:
--
通讯作者:
Eito Sato;Koji Inoue;Satoshi Kawakami
Eito Sato;Koji Inoue;Satoshi Kawakami
中科院分区:
其他
文献类型:
--
作者:
Eito Sato;Koji Inoue;Satoshi Kawakami

文献摘要

相似文献

循环神经网络 (RNN) 在处理时间序列数据的推理处理中取得了高性能。其中,用于快速处理 RNN 的硬件加速有助于实时性能至关重要的任务,例如语音识别和股市预测。纳米光子神经网络加速器是一种利用光的高速、高并行性和低功耗特性来实现高性能神经网络处理的方法。然而,由于缺乏递归路径和待设计模型的不成熟而导致巨大的开销,现有方法对于 RNN 来说效率低下。因此,利用 RNN 特性的架构考虑对于低延迟至关重要。本文提出了一种用于 RNN 的快速、低功耗处理单元,该单元引入了使用光学器件的激活函数和递归处理。我们阐明了噪声对所提出电路的计算精度和推理精度的影响。结果,计算精度随着递归次数的增加而显着下降,但对推理精度的影响可以忽略不计。我们还将所提出的电路的性能与全电设计和以光学方式处理矢量矩阵乘积并以电方式处理递归的混合设计进行了比较。因此,与全电气设计相比,该电路的性能将延迟提高了 467 倍,功耗降低了 93.0%;与混合设计相比,延迟提高了 7.3 倍,功耗降低了 58.6%。
Recurrent neural networks (RNNs) have achieved high performance in inference processing that handles time-series data. Among them, hardware acceleration for fast processing RNNs is helpful for tasks where real-time performance is es-sential, such as speech recognition and stock market prediction. The nano-photonic neural network accelerator is an approach that takes advantage of the high speed, high parallelism, and low power consumption of light to achieve high performance in neural network processing. However, existing methods are inefficient for RNNs due to significant overhead caused by the absence of recursive paths and the immaturity of the model to be designed. Therefore, architectural considerations that take advantage of RNN characteristics are essential for low latency. This paper proposes a fast and low-power processing unit for RNNs that introduces activation functions and recursion processing using optical devices. We clarified the impact of noise on the proposed circuit's calculation accuracy and inference accuracy. As a result, the calculation accuracy deteriorated significantly in proportion to the increase in the number of recursions, but the effect on inference accuracy was negligible. We also compared the performance of the proposed circuit to an all-electric design and a hybrid design that processes the vector-matrix product optically and the recursion electrically. As a result, the performance of the proposed circuit improves latency by 467x, reduces power consumption by 93.0% compared with the all-electrical design, improves latency by 7.3x, and reduces power consumption by 58.6% compared with the hybrid design.