Efficient Deep Neural Network Serving: Fast and Furious

Efficient Deep Neural Network Serving: Fast and Furious
复制标题

DOI:
10.1109/tnsm.2018.2808352
复制
发表时间:
2018-02
影响因子:
5.3
通讯作者:
Feng Yan;Yuxiong He;Olatunji Ruwase;E. Smirni
Feng Yan;Yuxiong He;Olatunji Ruwase;E. Smirni
中科院分区:
计算机科学2区
文献类型:
--
作者:
Feng Yan;Yuxiong He;Olatunji Ruwase;E. Smirni

文献摘要

被引文献

相似文献

深度神经网络(DNN)作为最先进的机器学习技术的出现,使图像识别、语音识别和翻译、药物发现和机器视觉等各种人工智能应用成为可能。这些应用程序由在云计算基础设施上以服务模式运行的大型DNN模型支持,以处理客户端输入,如图像,语音片段和文本片段。考虑到大型DNN模型的计算密集型性质,DNN服务系统的一个关键挑战是最小化请求响应延迟。本文描述了不同并行技术的行为,以支持大型DNN的可扩展和响应式服务系统。我们识别并建模DNN工作负载的两个重要属性:1)同构请求服务需求和2)由于缓存/内存争用而并发运行的请求之间的干扰。这些属性激发了快速服务深度学习系统(SERF)的设计,SERF是一种动态调度框架,由基于干扰感知的分析模型提供支持。为了最大限度地减少DNN服务的响应延迟,SERF通过使用经验和分析方法快速识别并切换到服务系统的最佳并行配置。我们使用几个著名的基准测试SERF的评估表明,它具有良好的延迟预测精度,能够正确识别每个基准测试的最佳并行配置,能够适应不断变化的负载条件,以及其效率优势(至少快三个数量级)。我们还证明了SERF支持其他调度目标,并可以扩展到任何一般的机器学习服务系统与类似的并行性能如上所述。
The emergence of deep neural networks (DNNs) as a state-of-the-art machine learning technique has enabled a variety of artificial intelligence applications for image recognition, speech recognition and translation, drug discovery, and machine vision. These applications are backed by large DNN models running in serving mode on a cloud computing infrastructure to process client inputs such as images, speech segments, and text segments. Given the compute-intensive nature of large DNN models, a key challenge for DNN serving systems is to minimize the request response latencies. This paper characterizes the behavior of different parallelism techniques for supporting scalable and responsive serving systems for large DNNs. We identify and model two important properties of DNN workloads: 1) homogeneous request service demand and 2) interference among requests running concurrently due to cache/memory contention. These properties motivate the design of serving deep learning systems fast (SERF), a dynamic scheduling framework that is powered by an interference-aware queueing-based analytical model. To minimize response latency for DNN serving, SERF quickly identifies and switches to the optimal parallel configuration of the serving system by using both empirical and analytical methods. Our evaluation of SERF using several well-known benchmarks demonstrates its good latency prediction accuracy, its ability to correctly identify optimal parallel configurations for each benchmark, its ability to adapt to changing load conditions, and its efficiency advantage (by at least three orders of magnitude faster) over exhaustive profiling. We also demonstrate that SERF supports other scheduling objectives and can be extended to any general machine learning serving system with the similar parallelism properties as above.