TMO: transparent memory offloading in datacenters

TMO: transparent memory offloading in datacenters
复制标题

DOI:
10.1145/3503222.3507731
复制
发表时间:
2022-02
期刊:
Proceedings of the 27th ACM International Conference on Architectural Support for Programming Languages and Operating Systems
影响因子:
--
通讯作者:
Johannes Weiner;Niket Agarwal;Dan Schatzberg;Leon Yang;Hao Wang-;Blaise Sanouillet;Bikash Sharma;Tejun Heo;M. Jain;Chunqiang Tang;Dimitrios Skarlatos
Johannes Weiner;Niket Agarwal;Dan Schatzberg;Leon Yang;Hao Wang-;Blaise Sanouillet;Bikash Sharma;Tejun Heo;M. Jain;Chunqiang Tang;Dimitrios Skarlatos
中科院分区:
其他
文献类型:
--
作者:
Johannes Weiner;Niket Agarwal;Dan Schatzberg;Leon Yang;Hao Wang-;Blaise Sanouillet;Bikash Sharma;Tejun Heo;M. Jain;Chunqiang Tang;Dimitrios Skarlatos

文献摘要

相似文献

新兴的数据中心应用的存储器需求的持续增长,沿着DRAM价格的不断增加的成本和波动性,已经导致DRAM成为主要的基础设施费用。替代技术,如NVMe SSD和即将推出的NVM设备,以成本和功耗的一小部分提供比DRAM更高的容量。一种有前途的方法是通过内核或管理程序技术将较冷的内存透明地卸载到较便宜的内存技术。然而,关键的挑战是开发一个企业级解决方案,该解决方案在处理不同的工作负载和不同卸载设备(如压缩内存、SSD和NVM)的大性能差异方面具有鲁棒性。本文介绍了TMO,Meta的异构数据中心环境的透明内存卸载解决方案。TMO引入了一种新的Linux内核机制,可以直接实时测量由于CPU、内存和I/O资源短缺而丢失的工作。在此信息的指导下,在没有任何先前应用知识的情况下,TMO自动调整要卸载到异构设备(例如,根据设备的性能特征和应用程序对存储器访问速度降低的敏感性,TMO不仅从应用程序容器,而且从提供基础设施级功能的sidecar容器全面识别卸载机会。为了最大限度地节省内存,TMO同时针对匿名内存和文件缓存,并平衡匿名内存的换入速率和最近从文件缓存中驱逐的文件页面的重新加载速率。TMO已经在生产环境中运行了一年多,在我们的大型数据中心机群中,已经为数百万台服务器节省了20-32%的总内存。我们已经成功地将TMO向上流到Linux内核中。
The unrelenting growth of the memory needs of emerging datacenter applications, along with ever increasing cost and volatility of DRAM prices, has led to DRAM being a major infrastructure expense. Alternative technologies, such as NVMe SSDs and upcoming NVM devices, offer higher capacity than DRAM at a fraction of the cost and power. One promising approach is to transparently offload colder memory to cheaper memory technologies via kernel or hypervisor techniques. The key challenge, however, is to develop a datacenter-scale solution that is robust in dealing with diverse workloads and large performance variance of different offload devices such as compressed memory, SSD, and NVM. This paper presents TMO, Meta’s transparent memory offloading solution for heterogeneous datacenter environments. TMO introduces a new Linux kernel mechanism that directly measures in realtime the lost work due to resource shortage across CPU, memory, and I/O. Guided by this information and without any prior application knowledge, TMO automatically adjusts how much memory to offload to heterogeneous devices (e.g., compressed memory or SSD) according to the device’s performance characteristics and the application’s sensitivity to memory-access slowdown. TMO holistically identifies offloading opportunities from not only the application containers but also the sidecar containers that provide infrastructure-level functions. To maximize memory savings, TMO targets both anonymous memory and file cache, and balances the swap-in rate of anonymous memory and the reload rate of file pages that were recently evicted from the file cache. TMO has been running in production for more than a year, and has saved between 20-32% of the total memory across millions of servers in our large datacenter fleet. We have successfully upstreamed TMO into the Linux kernel.