Reinforcement learning for resource management in multi-tenant serverless platforms

Reinforcement learning for resource management in multi-tenant serverless platforms
复制标题

DOI:
10.1145/3517207.3526971
复制
发表时间:
2022-04
期刊:
Proceedings of the 2nd European Workshop on Machine Learning and Systems
影响因子:
--
通讯作者:
Haoran Qiu;Weichao Mao;Archit Patke;Chen Wang;H. Franke;Z. Kalbarczyk;T. Başar;R. Iyer
Haoran Qiu;Weichao Mao;Archit Patke;Chen Wang;H. Franke;Z. Kalbarczyk;T. Başar;R. Iyer
中科院分区:
其他
文献类型:
--
作者:
Haoran Qiu;Weichao Mao;Archit Patke;Chen Wang;H. Franke;Z. Kalbarczyk;T. Başar;R. Iyer

文献摘要

被引文献

相似文献

无服务器功能即服务(FaaS)是一种新兴的云计算范例,它将应用程序开发人员从资源供应和扩展等基础设施管理任务中解放出来。为了减少函数的尾部延迟,提高资源利用率,近年来的研究重点是应用在线学习算法,如强化学习(RL)来管理资源。与现有的基于启发式的资源管理方法相比,基于强化学习的方法消除了人工参与,避免了启发式的产生。在本文中,我们展示了最先进的单代理RL算法(S-RL)在多租户无服务器FaaS平台上遭受高达4.6倍的功能尾部延迟退化,并且在训练期间无法收敛。然后,我们提出并实现了一种基于近端策略优化的定制多智能体RL算法,即多智能体PPO (MA-PPO)。我们表明,在多租户环境中,MA-PPO可以训练每个代理直到收敛,并提供与单租户情况下的S-RL相当的在线性能,并且性能下降不到10%。此外,在多租户情况下,MA-PPO提供了4.4倍的S-RL性能改进(就功能尾部延迟而言)。
Serverless Function-as-a-Service (FaaS) is an emerging cloud computing paradigm that frees application developers from infrastructure management tasks such as resource provisioning and scaling. To reduce the tail latency of functions and improve resource utilization, recent research has been focused on applying online learning algorithms such as reinforcement learning (RL) to manage resources. Compared to existing heuristics-based resource management approaches, RL-based approaches eliminate humans in the loop and avoid the painstaking generation of heuristics. In this paper, we show that the state-of-the-art single-agent RL algorithm (S-RL) suffers up to 4.6x higher function tail latency degradation on multi-tenant serverless FaaS platforms and is unable to converge during training. We then propose and implement a customized multi-agent RL algorithm based on Proximal Policy Optimization, i.e., multi-agent PPO (MA-PPO). We show that in multi-tenant environments, MA-PPO enables each agent to be trained until convergence and provides online performance comparable to S-RL in single-tenant cases with less than 10% degradation. Besides, MA-PPO provides a 4.4x improvement in S-RL performance (in terms of function tail latency) in multi-tenant cases.