nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devices

nn-Meter: towards accurate latency prediction of deep-learning model inference on diverse edge devices
复制标题

nn-Meter:在各种边缘设备上实现深度学习模型推理的准确延迟预测

DOI:
10.1145/3458864.3467882
复制
发表时间:
2021
期刊:
Proceedings of the 19th Annual International Conference on Mobile Systems, Applications, and Services
影响因子:
--
通讯作者:
Yunxin Liu
Yunxin Liu
中科院分区:
--
文献类型:
--
作者:
L. Zhang;S. Han;Jianyu Wei;Ningxin Zheng;Ting Cao;Yuqing Yang;Yunxin Liu

文献摘要

参考文献

被引文献

相似文献

随着设备端深度学习的最新趋势,推理延迟已成为在各种移动和边缘设备上运行深度神经网络(DNN)模型的一个关键指标。为此,在许多无法在真实设备上测量延迟或成本过高的任务中,DNN模型推理的延迟预测是非常必要的,例如从巨大的模型设计空间中搜索具有延迟约束的高效DNN模型。然而,这极具挑战性,并且由于不同边缘设备上的运行时优化导致模型推理延迟各不相同,现有方法无法实现高精度的预测。在本文中,我们提出并开发了nn - Meter,这是一种新颖且高效的系统,可准确预测DNN模型在不同边缘设备上的推理延迟。nn - Meter的关键思想是将整个模型推理划分为内核,即设备上的执行单元,并进行内核级别的预测。nn - Meter建立在两项关键技术之上:(i)内核检测,通过一组精心设计的测试用例自动检测模型推理的执行单元;(ii)自适应采样,从大空间中高效采样最有益的配置,以构建准确的内核级延迟预测器。在三种流行的边缘硬件平台(移动CPU、移动GPU和英特尔VPU)上实现,并使用包含26,000个模型的大型数据集进行评估,nn - Meter显著优于先前的最先进技术。
With the recent trend of on-device deep learning, inference latency has become a crucial metric in running Deep Neural Network (DNN) models on various mobile and edge devices. To this end, latency prediction of DNN model inference is highly desirable for many tasks where measuring the latency on real devices is infeasible or too costly, such as searching for efficient DNN models with latency constraints from a huge model-design space. Yet it is very challenging and existing approaches fail to achieve a high accuracy of prediction, due to the varying model-inference latency caused by the runtime optimizations on diverse edge devices. In this paper, we propose and develop nn-Meter, a novel and efficient system to accurately predict the inference latency of DNN models on diverse edge devices. The key idea of nn-Meter is dividing a whole model inference into kernels, i.e., the execution units on a device, and conducting kernel-level prediction. nn-Meter builds atop two key techniques: (i) kernel detection to automatically detect the execution unit of model inference via a set of well-designed test cases; and (ii) adaptive sampling to efficiently sample the most beneficial configurations from a large space to build accurate kernel-level latency predictors. Implemented on three popular platforms of edge hardware (mobile CPU, mobile GPU, and Intel VPU) and evaluated using a large dataset of 26,000 models, nn-Meter significantly outperforms the prior state-of-the-art.
DOI: 10.1145/3306346.3322967
发表时间: 2019-07-01
影响因子: 6.2
作者:
Adams, Andrew;Ma, Karima;Ragan-Kelley, Jonathan
通讯作者: Ragan-Kelley, Jonathan