LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks

LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks
复制标题

LaLaRAND:实时 DNN 任务的灵活的逐层 CPU/GPU 调度

DOI:
10.1109/rtss52674.2021.00038
复制
发表时间:
2021
期刊:
2021 IEEE Real-Time Systems Symposium (RTSS)
影响因子:
--
通讯作者:
H. Chwa
H. Chwa
中科院分区:
--
文献类型:
--
作者:
Woo;Kilho Lee;Jinkyu Lee;I. Shin;H. Chwa

文献摘要

参考文献

被引文献

相似文献

深度神经网络(DNN)在各种机器学习(ML)任务中取得了显着的成功,这些任务对许多安全关键的实时嵌入式系统都很有用。在实时嵌入式系统上实现DNN执行的首要设计目标是用有限的计算资源提供最坏情况下的时序保证。然而,最先进的ML框架几乎不利用异构计算资源(即,由于几个因素,包括粗粒度的资源分配模型(每个任务一个资源),DNN在CPU和GPU上执行的不对称性质,以及缺乏可扩展性感知的CPU/GPU分配方案,因此,DNN任务的可扩展性需要提高。本文介绍了,据我们所知,第一次研究解决上述三个主要障碍,并检查他们的合作效果,可兼容性的改善。在本文中,我们提出了LaLaRAND,一个实时层级DNN调度框架,通过将CPU友好量化与细粒度CPU/GPU分配方案(每层一个资源)紧密耦合,同时在不影响时序保证的情况下减轻准确性损失,从而实现各个DNN层的灵活CPU/GPU调度。我们已经在最先进的ML框架之上实现和评估了LaLaRAND,以证明其有效性,使更多的DNN任务集分别比现有方法和基线(vanilla PyTorch)高出56%和80%,性能(推理准确度)差异仅为-0.4%。
Deep neural networks (DNNs) have shown remarkable success in various machine-learning (ML) tasks useful for many safety-critical, real-time embedded systems. The foremost design goal for enabling DNN execution on real-time embedded systems is to provide worst-case timing guarantees with limited computing resources. Yet, the state-of-the-art ML frameworks hardly leverage heterogeneous computing resources (i.e., CPU, GPU) to improve the schedulability of real-time DNN tasks due to several factors, which include a coarse-grained resource allocation model (one-resource-per-task), the asymmetric nature of DNN execution on CPU and GPU, and lack of schedulability-aware CPU/GPU allocation scheme. This paper presents, to the best of our knowledge, the first study of addressing the above three major barriers and examining their cooperative effect on schedulability improvement. In this paper, we propose LaLaRAND, a real-time layer-level DNN scheduling framework, that enables flexible CPU/GPU scheduling of individual DNN layers by tightly coupling CPU-friendly quantization with fine-grained CPU/GPU allocation schemes (one-resource-per-layer) while mitigating accuracy loss without compromising timing guarantees. We have implemented and evaluated LaLaRAND on top of the state-of-the-art ML framework to demonstrate its effectiveness in making more DNN task sets schedulable by 56% and 80% over an existing approach and a baseline (vanilla PyTorch), respectively, with only up to -0.4% of performance (inference accuracy) difference.
DOI: 10.1145/3372224.3419209
发表时间: 2020-09
期刊: Proceedings of the 26th Annual International Conference on Mobile Computing and Networking
影响因子: --
作者:
Xingzhe Song;Boyuan Yang;Ge Yang;Ruirong Chen;E. Forno;Wei Chen;Wei Gao
通讯作者: Xingzhe Song;Boyuan Yang;Ge Yang;Ruirong Chen;E. Forno;Wei Chen;Wei Gao
重新思考时间敏感型自动驾驶应用的 CNN 框架:应对工业挑战
DOI: 10.1109/rtas.2019.00033
发表时间: 2019
期刊: 2019 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS
影响因子: --
作者:
Yang, Ming;Wang, Shige;Bakita, Joshua;Vu, Thanh;Smith, F. Donelson;Anderson, James H.;Frahm, Jan-Michael
通讯作者: Frahm, Jan-Michael