Polymorphic Accelerators for Deep Neural Networks

Polymorphic Accelerators for Deep Neural Networks
复制标题

深度神经网络的多态加速器

DOI:
10.1109/tc.2020.3048624
复制
发表时间:
2022-03
影响因子:
3.7
通讯作者:
Arash AziziMazreah;Lizhong Chen
Arash AziziMazreah;Lizhong Chen
中科院分区:
计算机科学2区
文献类型:
--
作者:
Arash AziziMazreah;Lizhong Chen

文献摘要

相似文献

深度神经网络(DNN)有多种形式,如卷积神经网络,多层感知器和递归神经网络,以满足机器学习应用的各种需求。然而,现有的DNN加速器设计在用于执行多个神经网络时,存在处理元件利用不足、特征图流量大以及面积开销大的问题。在这篇文章中,我们提出了一种新的方法,多态加速器,从根本上解决灵活性问题。我们引入逻辑加速器的抽象来解耦固定映射与物理资源。提出了三个程序,协同工作,重新配置加速器的当前网络正在执行,并使跨层的数据重用之间的逻辑加速器。评估结果表明,所提出的方法在数据重用,推理延迟和性能方面取得了显着改善,例如,与最先进的灵活带宽方法和资源划分方法相比,吞吐量分别增加了1.52倍和1.63倍。这证明了多态加速器架构的有效性和前景。
Deep neural networks (DNNs) come with many forms, such as convolutional neural networks, multilayer perceptron, and recurrent neural networks, to meet diverse needs of machine learning applications. However, existing DNN accelerator designs, when used to execute multiple neural networks, suffer from underutilization of processing elements, heavy feature map traffic, and large area overhead. In this article, we propose a novel approach, Polymorphic Accelerators, to address the flexibility issue fundamentally. We introduce the abstraction of logical accelerators to decouple the fixed mapping with physical resources. Three procedures are proposed that work collaboratively to reconfigure the accelerator for the current network that is being executed and to enable cross-layer data reuse among logical accelerators. Evaluation results show that the proposed approach achieves significant improvement in data reuse, inference latency and performance, e.g., 1.52x and 1.63x increase in throughput compared with state-of-the-art flexible dataflow approach and resource partitioning approach, respectively. This demonstrates the effectiveness and promise of polymorphic accelerator architecture.