Improving Performance Estimation for Design Space Exploration for Convolutional Neural Network Accelerators

Improving Performance Estimation for Design Space Exploration for Convolutional Neural Network Accelerators
复制标题

DOI:
10.3390/electronics10040520
复制
发表时间:
2021-02
期刊:
影响因子:
2.9
通讯作者:
Martin Ferianc;Hongxiang Fan;Divyansh Manocha;Hongyu Zhou;Shuanglong Liu;Xinyu Niu;W. Luk
Martin Ferianc;Hongxiang Fan;Divyansh Manocha;Hongyu Zhou;Shuanglong Liu;Xinyu Niu;W. Luk
中科院分区:
工程技术3区
文献类型:
--
作者:
Martin Ferianc;Hongxiang Fan;Divyansh Manocha;Hongyu Zhou;Shuanglong Liu;Xinyu Niu;W. Luk

文献摘要

被引文献

相似文献

神经网络(NN)的当代进展已经证明了它们在不同应用中的潜力,例如图像分类,对象检测或自然语言处理。特别是,可重构加速器已被广泛用于加速神经网络由于其可重构性和效率在特定的应用实例。为了确定加速器的配置,有必要进行设计空间探索以优化性能。然而,设计空间探索的过程是耗时的,因为缓慢的性能评估不同的配置。因此,需要一种准确且快速的性能预测方法来加速设计空间探索。这项工作介绍了一种新的方法,快速,准确地估计不同的指标,是重要的设计空间探索时。该方法是基于高斯过程回归模型参数化的加速器和目标NN的功能,以加速。我们评估了所提出的方法以及其他流行的基于机器学习的方法,以估计我们在针对卷积神经网络的两个不同硬件平台上实现的加速器的延迟和能耗。我们证明了估计精度的提高,而不需要显着的实施工作或调整。
Contemporary advances in neural networks (NNs) have demonstrated their potential in different applications such as in image classification, object detection or natural language processing. In particular, reconfigurable accelerators have been widely used for the acceleration of NNs due to their reconfigurability and efficiency in specific application instances. To determine the configuration of the accelerator, it is necessary to conduct design space exploration to optimize the performance. However, the process of design space exploration is time consuming because of the slow performance evaluation for different configurations. Therefore, there is a demand for an accurate and fast performance prediction method to speed up design space exploration. This work introduces a novel method for fast and accurate estimation of different metrics that are of importance when performing design space exploration. The method is based on a Gaussian process regression model parametrised by the features of the accelerator and the target NN to be accelerated. We evaluate the proposed method together with other popular machine learning based methods in estimating the latency and energy consumption of our implemented accelerator on two different hardware platforms targeting convolutional neural networks. We demonstrate improvements in estimation accuracy, without the need for significant implementation effort or tuning.