SIMPPO: a scalable and incremental online learning framework for serverless resource management

SIMPPO: a scalable and incremental online learning framework for serverless resource management
复制标题

DOI:
10.1145/3542929.3563475
复制
发表时间:
2022-11
期刊:
Proceedings of the 13th Symposium on Cloud Computing
影响因子:
--
通讯作者:
Haoran Qiu;Weichao Mao;Archit Patke;Chen Wang;H. Franke;Z. Kalbarczyk;T. Başar;R. Iyer
Haoran Qiu;Weichao Mao;Archit Patke;Chen Wang;H. Franke;Z. Kalbarczyk;T. Başar;R. Iyer
中科院分区:
其他
文献类型:
--
作者:
Haoran Qiu;Weichao Mao;Archit Patke;Chen Wang;H. Franke;Z. Kalbarczyk;T. Başar;R. Iyer

文献摘要

相似文献

无服务器功能即服务(FaaS)为客户提供了改进的可编程性,但它并不是“少”服务器,而是以云提供商更复杂的基础设施管理(例如,资源配置和调度)为代价。为了维持服务水平目标,提高资源利用效率,近年来的研究重点是应用在线学习算法(如强化学习(RL))来管理资源。尽管应用强化学习取得了初步成功,但我们首先在论文中表明,与孤立环境相比,最先进的单代理强化学习算法(S-RL)在多租户无服务器FaaS平台上的p99功能延迟降低高达4.8倍,并且在训练期间无法收敛。然后,我们设计并实现了一个基于近端策略优化(SIMPPO)的可扩展增量多智能体强化学习框架。我们的实验表明,在多租户环境中,SIMPPO使每个RL代理能够在训练期间有效地收敛,并提供与隔离训练的S-RL相当的在线功能延迟性能,并且性能下降很小(<9.2%)。此外,在多租户情况下,与S-RL相比,SIMPPO将p99功能延迟降低了4.5倍。
Serverless Function-as-a-Service (FaaS) offers improved programmability for customers, yet it is not server-"less" and comes at the cost of more complex infrastructure management (e.g., resource provisioning and scheduling) for cloud providers. To maintain service-level objectives (SLOs) and improve resource utilization efficiency, recent research has been focused on applying online learning algorithms such as reinforcement learning (RL) to manage resources. Despite the initial success of applying RL, we first show in this paper that the state-of-the-art single-agent RL algorithm (S-RL) suffers up to 4.8x higher p99 function latency degradation on multi-tenant serverless FaaS platforms compared to isolated environments and is unable to converge during training. We then design and implement a scalable and incremental multi-agent RL framework based on Proximal Policy Optimization (SIMPPO). Our experiments demonstrate that in multi-tenant environments, SIMPPO enables each RL agent to efficiently converge during training and provides online function latency performance comparable to that of S-RL trained in isolation with minor degradation (<9.2%). In addition, SIMPPO reduces the p99 function latency by 4.5x compared to S-RL in multi-tenant cases.