Hercules: Heterogeneity-Aware Inference Serving for At-Scale Personalized Recommendation

Hercules: Heterogeneity-Aware Inference Serving for At-Scale Personalized Recommendation
复制标题

DOI:
10.48550/arxiv.2203.07424
复制
发表时间:
2022-03
期刊:
2022 IEEE International Symposium on High-Performance Computer Architecture (HPCA)
影响因子:
--
通讯作者:
Liu Ke;Udit Gupta;Mark Hempstead;Carole-Jean Wu;Hsien-Hsin S. Lee;Xuan Zhang
Liu Ke;Udit Gupta;Mark Hempstead;Carole-Jean Wu;Hsien-Hsin S. Lee;Xuan Zhang
中科院分区:
其他
文献类型:
--
作者:
Liu Ke;Udit Gupta;Mark Hempstead;Carole-Jean Wu;Hsien-Hsin S. Lee;Xuan Zhang

文献摘要

被引文献

相似文献

个性化推荐是一类重要的深度学习应用程序,它支持大量的互联网服务,并消耗大量的数据中心资源。随着生产级推荐系统的规模不断增长,在异构数据中心中优化其服务性能和效率非常重要,并且可以转化为基础设施容量节省。在本文中,我们提出了Hercules,这是一个针对不同行业代表性模型和云规模异构系统的个性化推荐推理服务的优化框架。Hercules执行两个阶段的优化过程-离线分析和在线服务。第一阶段使用基于梯度的搜索算法搜索未开发的大型任务调度空间,在单个服务器上实现高达9.0倍的延迟限制吞吐量改进;它还为每个推荐工作负载确定最佳异构服务器架构。第二阶段执行异构感知集群配置,以优化资源映射和分配,以响应波动的昼夜负载。建议的集群调度器在大力神实现47.7%的集群容量节省,并减少了23.7%的最先进的贪婪调度器提供的功率。
Personalized recommendation is an important class of deep-learning applications that powers a large collection of internet services and consumes a considerable amount of datacenter resources. As the scale of production-grade recommendation systems continues to grow, optimizing their serving performance and efficiency in a heterogeneous datacenter is important and can translate into infrastructure capacity saving. In this paper, we propose Hercules, an optimized framework for personalized recommendation inference serving that targets diverse industry-representative models and cloud-scale heterogeneous systems. Hercules performs a two-stage optimization procedure — offline profiling and online serving. The first stage searches the large under-explored task scheduling space with a gradient-based search algorithm achieving up to 9.0× latency-bounded throughput improvement on individual servers; it also identifies the optimal heterogeneous server architecture for each recommendation workload. The second stage performs heterogeneity-aware cluster provisioning to optimize resource mapping and allocation in response to fluctuating diurnal loads. The proposed cluster scheduler in Hercules achieves 47.7% cluster capacity saving and reduces the provisioned power by 23.7% over a state-of-the-art greedy scheduler.