Automated Runtime-Aware Scheduling for Multi-Tenant DNN Inference on GPU

Automated Runtime-Aware Scheduling for Multi-Tenant DNN Inference on GPU
复制标题

DOI:
10.1109/iccad51958.2021.9643501
复制
发表时间:
2021-11
期刊:
2021 IEEE/ACM International Conference On Computer Aided Design (ICCAD)
影响因子:
--
通讯作者:
Fuxun Yu;Shawn Bray;Di Wang;Longfei Shangguan;Xulong Tang;Chenchen Liu;Xiang Chen
Fuxun Yu;Shawn Bray;Di Wang;Longfei Shangguan;Xulong Tang;Chenchen Liu;Xiang Chen
中科院分区:
其他
文献类型:
--
作者:
Fuxun Yu;Shawn Bray;Di Wang;Longfei Shangguan;Xulong Tang;Chenchen Liu;Xiang Chen

文献摘要

相似文献

随着深度神经网络(DNN)的快速发展,许多现实世界的应用正在采用多个模型来执行复合任务,例如在自动驾驶汽车上共同运行分类,检测和分割模型。这种多租户DNN推理案例大大加剧了计算复杂性,并要求对图级操作员调度、运行时级资源感知以及硬件调度器支持进行全面协作。然而,目前对这种多租户推理的调度支持仍然相对落后。在这项工作中,我们提出了一个资源感知的调度框架,用于在GPU上进行高效的多租户DNN推理,该框架自动协调不同执行级别的DNN计算。利用统一的调度中间表示和自动化的基于ML的搜索算法,可以生成最佳调度,以明智地调整模型并发性和交错DNN模型运算符,在整个推理过程中保持持续平衡的资源利用率,并最终提高运行时效率。实验表明,与常规DNN运行时库(例如,CuDNN、TVM)和特定的并发调度方法(例如,NVIDIA Multi-Stream)。
With the fast development of deep neural networks (DNNs), many real-world applications are adopting multiple models to conduct compound tasks, such as co-running classification, detection, and segmentation models on autonomous vehicles. Such multi-tenant DNN inference cases greatly exacerbate the computational complexity and call for comprehensive collaboration for graph-level operator scheduling, runtime-level resource awareness, as well as hardware scheduler support. However, the current scheduling support for such multi-tenant inference is still relatively backward. In this work, we propose a resource-aware scheduling framework for efficient multi-tenant DNN inference on GPU, which automatically coordinates DNN computing in different execution levels. Leveraging the unified scheduling intermediate representation and the automated ML-based searching algorithm, optimal schedules could be generated to wisely adjust model concurrency and interleave DNN model operators, maintaining a continuously balanced resource utilization across the entire inference process, and eventually improving the runtime efficiency. Experiments show that we could consistently achieve $1.3\times\sim 1.7\times$ speed-up, comparing to regular DNN runtime libraries (e.g., CuDNN, TVM) and particular concurrent scheduling methods (e.g., NVIDIA Multi-Stream).