Scheduling Real-time Deep Learning Services as Imprecise Computations

Scheduling Real-time Deep Learning Services as Imprecise Computations
复制标题

DOI:
10.1109/rtcsa50079.2020.9203676
复制
发表时间:
2020-08
期刊:
2020 IEEE 26th International Conference on Embedded and Real-Time Computing Systems and Applications (RTCSA)
影响因子:
--
通讯作者:
Shuochao Yao;Yifan Hao;Yiran Zhao;Huajie Shao;Dongxin Liu;Shengzhong Liu;Tianshi Wang;Jinyang Li;T. Abdelzaher
Shuochao Yao;Yifan Hao;Yiran Zhao;Huajie Shao;Dongxin Liu;Shengzhong Liu;Tianshi Wang;Jinyang Li;T. Abdelzaher
中科院分区:
其他
文献类型:
--
作者:
Shuochao Yao;Yifan Hao;Yiran Zhao;Huajie Shao;Dongxin Liu;Shengzhong Liu;Tianshi Wang;Jinyang Li;T. Abdelzaher

文献摘要

被引文献

相似文献

该论文提出了一种用于智能实时边缘服务的实时计算框架,代表本身无法支持大量计算的本地嵌入式设备。这项工作为实时计算开辟了一个新方向,为机器智能任务开发调度算法,以实现随时预测。我们证明深度神经网络工作流程可以被转换为不精确的计算,每个计算都有一个强制性部分和(几个)可选部分,其执行效用取决于输入数据。通过我们的设计,深度神经网络可以在完成之前被抢占,并支持随时推理。实时调度程序的目标是最大限度地提高深度神经网络输出的平均准确度,同时满足任务期限,这要归功于机会性地放弃最少必要的可选部分。这项工作的动机是日益普遍但资源有限的嵌入式设备(适用于从自动驾驶汽车到物联网的应用)的激增以及开发赋予它们智能的服务的愿望。在最新 GPU 硬件和最先进的机器视觉深度神经网络上进行的实验表明,我们的方案可以将整体精度提高 10% ∼ 20%,同时(几乎)不会错过最后期限。
The paper presents a real-time computing framework for intelligent real-time edge services, on behalf of local embedded devices that are themselves unable to support extensive computations. The work contributes to a new direction in realtime computing that develops scheduling algorithms for machine intelligence tasks that enable anytime prediction. We show that deep neural network workflows can be cast as imprecise computations, each with a mandatory part and (several) optional parts whose execution utility depends on input data. With our design, deep neural networks can be preempted before their completion and support anytime inference. The goal of the realtime scheduler is to maximize the average accuracy of deep neural network outputs while meeting task deadlines, thanks to opportunistic shedding of the least necessary optional parts. The work is motivated by the proliferation of increasingly ubiquitous but resource-constrained embedded devices (for applications ranging from autonomous cars to the Internet of Things) and the desire to develop services that endow them with intelligence. Experiments on recent GPU hardware and a state of the art deep neural network for machine vision illustrate that our scheme can increase the overall accuracy by 10% ∼ 20% while incurring (nearly) no deadline misses.