S^3DNN: Supervised Streaming and Scheduling for GPU-Accelerated Real-Time DNN Workloads

S^3DNN: Supervised Streaming and Scheduling for GPU-Accelerated Real-Time DNN Workloads
复制标题

S^3D​​NN:GPU 加速的实时 DNN 工作负载的监督流和调度

DOI:
--
复制
发表时间:
2018
期刊:
IEEE Real Time Technology and Applications Symposium
影响因子:
--
通讯作者:
Cong Liu
Cong Liu
中科院分区:
--
文献类型:
--
作者:
Husheng Zhou;Soroush Bateni;Cong Liu

文献摘要

被引文献

相似文献

深度神经网络(DNN)被广泛应用于许多需要自主决策的先进嵌入式系统中,例如自动驾驶和机器人。为了处理对资源要求很高的DNN工作负载,图形处理器(GPU)被用作主要的加速引擎。虽然已经进行了大量的研究来从算法上优化将DNN应用于目标识别等应用的效率,但在系统级优化GPU加速的DNN工作负载的执行方面的关注有限。本文提出了一种在实时多任务环境中优化动态神经网络负载在图形处理器上执行的系统解决方案S^3DNN,它同时优化了实时正确性和吞吐量这两个(有时是冲突的)目标。S^3DNN包含一个调控器,它选择性地收集系统范围的DNN请求以执行智能数据融合,以及一个新颖的监督流和调度框架,该框架将截止日期感知调度器与启用并发的CUDA流技术相结合。为了同时最大化并发带来的收益和实时性能,S^3DNN探索了分布式神经网络负载的一个相当有趣和独特的特征,其中一个分布式神经网络实例的多层经常呈现出逐渐降低的GPU资源利用率模式。我们已经在图形处理器加速的系统中全面实现了S^3DNN,并进行了大量的实验,评估了S^3DNN在各种系统和工作负载场景下的性能。测试结果表明,S^3DNN在实时性能和吞吐量方面较现有的GPU加速DNN处理框架有了显著的提高,分别提高了37%和40%以上。
Deep Neural Networks (DNNs) are being widely applied in many advanced embedded systems that require autonomous decision making, e.g., autonomous driving and robotics. To handle resource-demanding DNN workloads, graphic processing units (GPUs) have been used as the main acceleration engine. Although much research has been conducted to algorithmically optimize the efficiency of applying DNN to applications such as object recognition, limited attention has been given to optimizing the execution of GPU-accelerated DNN workloads at the system level. In this paper, we propose S^3DNN, a system solution that optimizes the execution of DNN workloads on GPU in a real-time multi-tasking environment, which simultaneously optimizes the two (sometimes) conflicting goals of real-time correctness and throughput. S^3DNN contains a governor that selectively gathers system-wide DNN requests to perform smart data fusion, and a novel supervised streaming and scheduling framework that combines a deadline-aware scheduler with the concurrency-enabled CUDA stream technique. To simultaneously maximize concurrency-induced benefits and real-time performance, S^3DNN explores a rather interesting and unique characteristic of DNN workloads, where multiple layers of a DNN instance often exhibit a gradually decreased GPU resource utilization pattern. We have fully implemented S^3DNN in a GPU-accelerated system and have conducted extensive sets of experiments evaluating the efficacy of S^3DNN under a wide range of system and workload scenarios. The results show that S^3DNN significantly improves upon state-of-the-art GPU-accelerated DNN processing frameworks, e.g., up to 37% and over 40% improvements in real-time performance and throughput, respectively.