A Machine Learning Framework to Improve Storage System Performance

A Machine Learning Framework to Improve Storage System Performance
复制标题

DOI:
10.1145/3465332.3470875
复制
发表时间:
2021-07
期刊:
Proceedings of the 13th ACM Workshop on Hot Topics in Storage and File Systems
影响因子:
--
通讯作者:
I. Akgun;A. S. Aydin;Aadil Shaikh;L. Velikov;E. Zadok
I. Akgun;A. S. Aydin;Aadil Shaikh;L. Velikov;E. Zadok
中科院分区:
其他
文献类型:
--
作者:
I. Akgun;A. S. Aydin;Aadil Shaikh;L. Velikov;E. Zadok

文献摘要

被引文献

相似文献

存储系统及其操作系统组件旨在适应各种应用程序和动态工作负载。操作系统内部的存储组件包含各种启发式算法,为不同的工作负载提供高性能和适应性。可以通过参数对这些方法进行调整,并且某些系统调用允许用户优化其系统性能。这些参数通常是基于有限应用和硬件的实验而预先确定的。因此,存储系统经常以这些预定的并且可能是次优的值运行。手动调整这些参数是不切实际的:人们需要一个自适应的智能系统来处理动态和复杂的工作负载。机器学习(ML)技术能够识别模式,对其进行抽象,并对新数据进行预测。ML可以成为优化和适应存储系统的关键组件。在这份立场文件中,我们提出了KML,这是一个用于存储系统的ML框架。我们实现了一个原型,并证明了它的能力,众所周知的问题,调整最佳的预读值。我们的研究结果表明,KML的内存占用量很小,引入的开销可以忽略不计,但吞吐量却提高了2.3倍。
Storage systems and their OS components are designed to accommodate a wide variety of applications and dynamic workloads. Storage components inside the OS contain various heuristic algorithms to provide high performance and adaptability for different workloads. These heuristics may be tunable via parameters, and some system calls allow users to optimize their system performance. These parameters are often predetermined based on experiments with limited applications and hardware. Thus, storage systems often run with these predetermined and possibly suboptimal values. Tuning these parameters manually is impractical: one needs an adaptive, intelligent system to handle dynamic and complex workloads. Machine learning (ML) techniques are capable of recognizing patterns, abstracting them, and making predictions on new data. ML can be a key component to optimize and adapt storage systems. In this position paper, we propose KML, an ML framework for storage systems. We implemented a prototype and demonstrated its capabilities on the well-known problem of tuning optimal readahead values. Our results show that KML has a small memory footprint, introduces negligible overhead, and yet enhances throughput by as much as 2.3x.