OneAdapt: Fast Adaptation for Deep Learning Applications via Backpropagation

OneAdapt: Fast Adaptation for Deep Learning Applications via Backpropagation
复制标题

DOI:
10.1145/3620678.3624653
复制
发表时间:
2023-10
期刊:
Proceedings of the 2023 ACM Symposium on Cloud Computing
影响因子:
--
通讯作者:
Kuntai Du;Yuhan Liu;Yitian Hao;Qizheng Zhang;Haodong Wang;Yuyang Huang;Ganesh Ananthanarayanan;Junchen Jiang
Kuntai Du;Yuhan Liu;Yitian Hao;Qizheng Zhang;Haodong Wang;Yuyang Huang;Ganesh Ananthanarayanan;Junchen Jiang
中科院分区:
其他
文献类型:
--
作者:
Kuntai Du;Yuhan Liu;Yitian Hao;Qizheng Zhang;Haodong Wang;Yuyang Huang;Ganesh Ananthanarayanan;Junchen Jiang

文献摘要

相似文献

流媒体数据的深度学习推理,如视频或激光雷达馈电中的目标检测以及音频波中的文本提取,现在无处不在。为了达到较高的推理精度,这些应用通常需要大量的网络带宽来收集高保真数据和大量的GPU资源来运行深度神经网络(dnn)。虽然对网络带宽和GPU资源的高需求可以通过优化调整配置旋钮(如视频分辨率和帧速率)来大幅降低,但目前的适应技术无法同时满足三个要求:调整配置(i)以最小的额外GPU或带宽开销(ii)根据数据如何影响最终DNN的准确性达到近乎最佳的决策,以及(iii)对一系列配置旋钮进行调整。本文介绍了OneAdapt,它通过利用梯度上升策略来适应配置旋钮来满足这些需求。关键思想是利用dnn的可微分性来快速估计每个配置旋钮的精度梯度,称为AccGrad。具体来说,OneAdapt通过乘以两个梯度来估计AccGrad: InputGrad(即,每个配置旋钮如何影响DNN的输入)和DNNGrad(即,DNN输入如何影响DNN推理输出)。我们通过五种类型的配置、四种分析任务和五种类型的输入数据来评估OneAdapt。与最先进的适配方案相比,OneAdapt在使用相同或更少的资源的情况下,在保持相当精度的同时,将带宽使用和GPU使用减少了15-59%,或将精度提高了1-5%。
Deep learning inference on streaming media data, such as object detection in video or LiDAR feeds and text extraction from audio waves, is now ubiquitous. To achieve high inference accuracy, these applications typically require significant network bandwidth to gather high-fidelity data and extensive GPU resources to run deep neural networks (DNNs). While the high demand for network bandwidth and GPU resources could be substantially reduced by optimally adapting the configuration knobs, such as video resolution and frame rate, current adaptation techniques fail to meet three requirements simultaneously: adapt configurations (i) with minimum extra GPU or bandwidth overhead (ii) to reach near-optimal decisions based on how the data affects the final DNN's accuracy, and (iii) do so for a range of configuration knobs. This paper presents OneAdapt, which meets these requirements by leveraging a gradient-ascent strategy to adapt configuration knobs. The key idea is to embrace DNNs' differentiability to quickly estimate the accuracy's gradient to each configuration knob, called AccGrad. Specifically, OneAdapt estimates AccGrad by multiplying two gradients: InputGrad (i.e., how each configuration knob affects the input to the DNN) and DNNGrad (i.e., how the DNN input affects the DNN inference output). We evaluate OneAdapt across five types of configurations, four analytic tasks, and five types of input data. Compared to state-of-the-art adaptation schemes, OneAdapt cuts bandwidth usage and GPU usage by 15-59% while maintaining comparable accuracy or improves accuracy by 1-5% while using equal or fewer resources.