Trimmer: Cost-Efficient Deep Learning Auto-tuning for Cloud Datacenters

Trimmer: Cost-Efficient Deep Learning Auto-tuning for Cloud Datacenters
复制标题

DOI:
10.1109/cloud55607.2022.00061
复制
发表时间:
2022-07
期刊:
2022 IEEE 15th International Conference on Cloud Computing (CLOUD)
影响因子:
--
通讯作者:
Damian Borowiec;Gingfung Yeung;A. Friday;Richard Harper;Peter Garraghan
Damian Borowiec;Gingfung Yeung;A. Friday;Richard Harper;Peter Garraghan
中科院分区:
其他
文献类型:
--
作者:
Damian Borowiec;Gingfung Yeung;A. Friday;Richard Harper;Peter Garraghan

文献摘要

相似文献

能够以更低的资源成本提供高性能机器学习即服务(MLaaS)的云计算中心是通过自动调整实现的:深度学习模型的自动张量程序优化,以最大限度地减少硬件设备内的推理延迟。然而,考虑到深度学习模型、库和硬件设备的广泛异构性,在云计算中心内执行自动调优会导致大量的时间、计算资源和能源成本,而最先进的自动调优并不能减轻这些成本。在本文中,我们提出了Trimmer,这是一个高性能和高成本效益的深度学习自动调优框架,用于云计算中心。Trimmer通过抢占表现出较差优化改进的张量程序实现来最大化DL模型性能和张量程序成本效率;并应用基于ML的过滤方法来替换昂贵的低性能张量程序,以提供选择低延迟张量程序的更大可能性。通过探索DL模型优化技术成本的实证研究,我们的分析表明,总能量的26-43%花费在测量张量程序实现上,这些实现对自动调整没有积极贡献。实验结果表明,Trimmer在不同的DL模型之间实现了高的自动调整成本效率,并将云集群的自动调整能耗降低了21.8-40.9%,同时实现了与最先进技术相当的DL模型延迟。
Cloud datacenters capable of provisioning high performance Machine Learning-as-a-Service (MLaaS) at reduced resource cost is achieved via auto-tuning: automated tensor program optimization of Deep Learning models to minimize inference latency within a hardware device. However given the extensive heterogeneity of Deep Learning models, libraries, and hardware devices, performing auto-tuning within Cloud datacenters incurs a significant time, compute resource, and energy cost of which state-of-the-art auto-tuning is not designed to mitigate. In this paper we propose Trimmer, a high performance and cost-efficient Deep Learning auto-tuning framework for Cloud datacenters. Trimmer maximizes DL model performance and tensor program cost-efficiency by preempting tensor program implementations exhibiting poor optimization improvement; and applying an ML-based filtering method to replace expensive low performing tensor programs to provide greater likelihood of selecting low latency tensor programs. Through an empirical study exploring the cost of DL model optimization techniques, our analysis indicates that 26–43% of total energy is expended on measuring tensor program implementations that do not positively contribute towards auto-tuning. Experiment results show that Trimmer achieves high auto-tuning cost-efficiency across different DL models, and reduces auto-tuning energy use by 21.8–40.9% for Cloud clusters whilst achieving DL model latency equivalent to state-of-the-art techniques.