Improving Storage Systems Using Machine Learning

Improving Storage Systems Using Machine Learning
复制标题

DOI:
10.1145/3568429
复制
发表时间:
2022-11
影响因子:
1.7
通讯作者:
I. Akgun;A. S. Aydin;Andrew Burford;Michael McNeill;Michael Arkhangelskiy;E. Zadok
I. Akgun;A. S. Aydin;Andrew Burford;Michael McNeill;Michael Arkhangelskiy;E. Zadok
中科院分区:
计算机科学3区
文献类型:
--
作者:
I. Akgun;A. S. Aydin;Andrew Burford;Michael McNeill;Michael Arkhangelskiy;E. Zadok

文献摘要

相似文献

操作系统包括许多旨在提高整体存储性能和吞吐量的启发式算法。由于这种启发式方法不能很好地适用于所有条件和工作负载,因此系统设计人员求助于向用户提供大量可调参数,从而加重用户不断优化自己的存储系统和应用程序的负担。存储系统通常负责I/O密集型应用程序中的大部分延迟,因此即使是很小的延迟改进也可能是显著的。机器学习(ML)技术承诺学习模式,从中归纳,并支持适应不断变化的工作负载的最佳解决方案。我们建议ML解决方案成为OSS中的一流组件,并取代人工启发式方法来动态优化存储系统。在本文中,我们描述了我们提出的称为KML的ML体系结构。我们开发了一个原型KML架构,并将其应用于两个案例研究:优化预读取值和NFS读取大小值。我们的实验表明,KML消耗不到4KB的动态内核内存,CPU开销小于0.2%,但对于两个案例研究-即使是在不同存储设备上同时运行的复杂、前所未见的混合工作负载,KML也可以学习模式并将I/O吞吐量提高高达2.3倍和15倍。
Operating systems include many heuristic algorithms designed to improve overall storage performance and throughput. Because such heuristics cannot work well for all conditions and workloads, system designers resorted to exposing numerous tunable parameters to users—thus burdening users with continually optimizing their own storage systems and applications. Storage systems are usually responsible for most latency in I/O-heavy applications, so even a small latency improvement can be significant. Machine learning (ML) techniques promise to learn patterns, generalize from them, and enable optimal solutions that adapt to changing workloads. We propose that ML solutions become a first-class component in OSs and replace manual heuristics to optimize storage systems dynamically. In this article, we describe our proposed ML architecture, called KML. We developed a prototype KML architecture and applied it to two case studies: optimizing readahead and NFS read-size values. Our experiments show that KML consumes less than 4 KB of dynamic kernel memory, has a CPU overhead smaller than 0.2%, and yet can learn patterns and improve I/O throughput by as much as 2.3× and 15× for two case studies—even for complex, never-seen-before, concurrently running mixed workloads on different storage devices.