Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and Scheduling

Accelerating End-to-End Deep Learning Workflow With Codesign of Data Preprocessing and Scheduling
复制标题

通过数据预处理和调度的协同设计加速端到端深度学习工作流程

DOI:
10.1109/tpds.2020.3047966
复制
发表时间:
2021-07-01
影响因子:
5.3
通讯作者:
Xiong, Yongqiang
Xiong, Yongqiang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Cheng, Yang;Li, Dan;Xiong, Yongqiang

文献摘要

被引文献

相似文献

在本文中,我们研究了现有深度学习(DL)系统的性能瓶颈,并提出了DLBooster来提高在GPU集群上部署DL应用程序的运行效率。在其核心,DLBooster利用两级优化来提升端到端DL工作流程。一方面,DLBooster选择性地将一些关键的解码工作负载卸载到FPGA,为计算引擎提供高性能的在线数据预处理服务。另一方面,DLBooster使用反向传播算法重新组织训练神经网络的计算工作量,并根据它们的依赖关系对其进行调度,以提高运行时GPU的利用率。基于我们的实验,我们表明,与基线相比,DLBooster可以提高图像处理吞吐量1.4× - 2.5×,并减少处理延迟的1/3,在几个真实世界的DL应用程序和数据集。此外,DLBooster在运行时管理FPGA设备所需的CPU核心不到1个,在某些情况下比基线至少少90%。DLBooster显示了其在云中加速DL工作流的潜力。
In this article, we investigate the performance bottleneck of existing deep learning (DL) systems and propose DLBooster to improve the running efficiency of deploying DL applications on GPU clusters. At its core, DLBooster leverages two-level optimizations to boost the end-to-end DL workflow. On the one hand, DLBooster selectively offloads some key decoding workloads to FPGAs to provide high-performance online data preprocessing services to the computing engine. On the other hand, DLBooster reorganizes the computational workloads of training neural networks with the backpropagation algorithm and schedules them according to their dependencies to improve the utilization of GPUs at runtime. Based on our experiments, we demonstrate that compared with baselines, DLBooster can improve the image processing throughput by 1.4× – 2.5× and reduce the processing latency by 1/3 in several real-world DL applications and datasets. Moreover, DLBooster consumes less than 1 CPU core to manage FPGA devices at runtime, which is at least 90 percent less than the baselines in some cases. DLBooster shows its potential to accelerate DL workflows in the cloud.