Mind the Gap: Broken Promises of CPU Reservations in Containerized Multi-tenant Clouds

Mind the Gap: Broken Promises of CPU Reservations in Containerized Multi-tenant Clouds
复制标题

DOI:
10.1145/3472883.3486997
复制
发表时间:
2021-11
期刊:
Proceedings of the ACM Symposium on Cloud Computing
影响因子:
--
通讯作者:
Li Liu;Haoliang Wang;An Wang;Mengbai Xiao;Yue Cheng;Songqing Chen
Li Liu;Haoliang Wang;An Wang;Mengbai Xiao;Yue Cheng;Songqing Chen
中科院分区:
其他
文献类型:
--
作者:
Li Liu;Haoliang Wang;An Wang;Mengbai Xiao;Yue Cheng;Songqing Chen

文献摘要

被引文献

相似文献

容器化正变得越来越流行,但不幸的是,容器经常无法使用分配的资源交付预期的性能。在本文中,我们首先演示了在容器共存的多租户环境中,性能差异和性能下降是显著的(高达5倍)。然后,我们研究这种性能下降的根本原因。与通常认为这种退化是由资源争用和干扰引起的观点相反,我们发现容器预留的CPU数量与实际获得的CPU数量之间存在差距。根本原因在于当今Linux调度机制的设计选择,我们称之为强制运行队列共享和虚拟CPU时间。实际上,预留CPU资源的需求与完全公平调度程序的工作保护特性之间存在着根本的冲突,这种矛盾使容器无法充分利用其请求的CPU资源。作为概念验证,我们在广泛使用的Kubernetes和Linux上实现了一种新的资源配置机制,以展示其潜在的好处,并为未来重新设计调度器提供思路。与现有的调度程序相比,我们的概念验证将批处理和交互式容器化应用程序的性能分别提高了5.6倍和13.7倍。
Containerization is becoming increasingly popular, but unfortunately, containers often fail to deliver the anticipated performance with the allocated resources. In this paper, we first demonstrate the performance variance and degradation are significant (by up to 5x) in a multi-tenant environment where containers are co-located. We then investigate the root cause of such performance degradation. Contrary to the common belief that such degradation is caused by resource contention and interference, we find that there is a gap between the amount of CPU a container reserves and actually gets. The root cause lies in the design choices of today's Linux scheduling mechanism, which we call Forced Runqueue Sharing and Phantom CPU Time. In fact, there are fundamental conflicts between the need to reserve CPU resources and Completely Fair Scheduler's work-conserving nature, and this contradiction prevents a container from fully utilizing its requested CPU resources. As a proof-of-concept, we implement a new resource configuration mechanism atop the widely used Kubernetes and Linux to demonstrate its potential benefits and shed light on future scheduler redesign. Our proof-of-concept, compared to the existing scheduler, improves the performance of both batch and interactive containerized apps by up to 5.6x and 13.7x.