LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks
LaLaRAND: Flexible Layer-by-Layer CPU/GPU Scheduling for Real-Time DNN Tasks
复制标题
LaLaRAND:实时 DNN 任务的灵活的逐层 CPU/GPU 调度
DOI:
10.1109/rtss52674.2021.00038
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
H. Chwa
中科院分区:
文献类型:
--
作者:
Woo;Kilho Lee;Jinkyu Lee;I. Shin;H. Chwa
Deep neural networks (DNNs) have shown remarkable success in various machine-learning (ML) tasks useful for many safety-critical, real-time embedded systems. The foremost design goal for enabling DNN execution on real-time embedded systems is to provide worst-case timing guarantees with limited computing resources. Yet, the state-of-the-art ML frameworks hardly leverage heterogeneous computing resources (i.e., CPU, GPU) to improve the schedulability of real-time DNN tasks due to several factors, which include a coarse-grained resource allocation model (one-resource-per-task), the asymmetric nature of DNN execution on CPU and GPU, and lack of schedulability-aware CPU/GPU allocation scheme. This paper presents, to the best of our knowledge, the first study of addressing the above three major barriers and examining their cooperative effect on schedulability improvement. In this paper, we propose LaLaRAND, a real-time layer-level DNN scheduling framework, that enables flexible CPU/GPU scheduling of individual DNN layers by tightly coupling CPU-friendly quantization with fine-grained CPU/GPU allocation schemes (one-resource-per-layer) while mitigating accuracy loss without compromising timing guarantees. We have implemented and evaluated LaLaRAND on top of the state-of-the-art ML framework to demonstrate its effectiveness in making more DNN task sets schedulable by 56% and 80% over an existing approach and a baseline (vanilla PyTorch), respectively, with only up to -0.4% of performance (inference accuracy) difference.
DOI:
10.1145/3372224.3419209
发表时间:
2020-09
期刊:
Proceedings of the 26th Annual International Conference on Mobile Computing and Networking
影响因子:
--
作者:
Xingzhe Song;Boyuan Yang;Ge Yang;Ruirong Chen;E. Forno;Wei Chen;Wei Gao
通讯作者:
Xingzhe Song;Boyuan Yang;Ge Yang;Ruirong Chen;E. Forno;Wei Chen;Wei Gao
DOI:
10.1109/rtas.2019.00033
发表时间:
2019
期刊:
2019 IEEE Real-Time and Embedded Technology and Applications Symposium (RTAS
影响因子:
--
作者:
Yang, Ming;Wang, Shige;Bakita, Joshua;Vu, Thanh;Smith, F. Donelson;Anderson, James H.;Frahm, Jan-Michael
通讯作者:
Frahm, Jan-Michael