Carver: Finding Important Parameters for Storage System Tuning

Carver: Finding Important Parameters for Storage System Tuning
复制标题

DOI:
--
复制
发表时间:
2020
影响因子:
42.7
通讯作者:
Zhen Cao;G. Kuenning;E. Zadok
Zhen Cao;G. Kuenning;E. Zadok
中科院分区:
生物学1区
文献类型:
--
作者:
Zhen Cao;G. Kuenning;E. Zadok

文献摘要

被引文献

相似文献

存储系统通常有许多影响其行为的参数。调优这些参数可以显著提高性能。唉,手动和自动调优方法由于大量的参数和指数数量的可能配置而陷入困境。由于先前的研究表明,一些参数比其他参数对性能的影响更大,因此专注于少数更重要的参数可以加快自动调优系统的速度,因为它们将具有更小的状态空间来探索。在本文中,我们提出了Carver,它使用(1)基于方差的度量来量化存储参数的重要性,(2)拉丁超立方体采样(Latin Hypercube Sampling)对巨大的参数空间进行采样;(3)一种贪婪而高效的参数选择算法,能够识别出重要的参数。我们在数据集上对Carver进行了评估,这些数据集由7个文件系统上的50多万个实验组成,包含4种代表性工作负载。Carver成功地确定了所有文件系统的重要参数,并表明其重要性随工作负载的不同而不同。我们证明了Carver能够在我们的数据集中识别出一组近乎最优的重要参数。我们用一小部分数据集测试了Carver的效率;它能够仅用整个数据集的0.4%就识别出相同的一组重要参数。
Storage systems usually have many parameters that affect their behavior. Tuning those parameters can provide significant gains in performance. Alas, both manual and automatic tuning methods struggle due to the large number of parameters and exponential number of possible configu-rations. Since previous research has shown that some parameters have greater performance impact than others, focusing on a smaller number of more important parameters can speed up auto-tuning systems because they would have a smaller state space to explore. In this paper, we propose Carver, which uses (1) a variance-based metric to quantify storage parameters’ importance, (2) Latin Hypercube Sampling to sample huge parameter spaces; and (3) a greedy but efficient parameter-selection algorithm that can identify important parameters. We evaluated Carver on datasets consisting of more than 500,000 experiments on 7 file systems, under 4 representative workloads. Carver successfully iden-tified important parameters for all file systems and showed that importance varies with different workloads. We demonstrated that Carver was able to identify a near-optimal set of important parameters in our datasets. We showed Carver’s efficiency by testing it with a small fraction of our dataset; it was able to identify the same set of important parameters with as little as 0.4% of the whole dataset.