Guided Bayesian Optimization to AutoTune Memory-Based Analytics

Guided Bayesian Optimization to AutoTune Memory-Based Analytics
复制标题

DOI:
10.1109/icdew.2019.00-22
复制
发表时间:
2019-04
期刊:
2019 IEEE 35th International Conference on Data Engineering Workshops (ICDEW)
影响因子:
--
通讯作者:
Mayuresh Kunjir
Mayuresh Kunjir
中科院分区:
其他
文献类型:
--
作者:
Mayuresh Kunjir

文献摘要

被引文献

相似文献

今天,人们对构建自主(或自动驾驶)数据处理系统很感兴趣。一个新兴的思想流派是利用贝叶斯优化的“黑箱”算法来解决这种问题,这是因为它具有更广泛的适用性和对结果质量的理论保证。然而,黑箱方法可能是时间和劳动密集型的;否则就会陷入局部最小值。我们研究了一个重要的问题,自动调整的内存分配的应用程序运行在现代分布式数据处理系统。一个简单的“白盒”模型的开发,可以快速区分好的配置和坏的。为了将这两种调优方法的优点联合收割机结合起来,我们构建了一个名为Guided Bayesian Optimization(GBO)的框架,该框架在Bayesian Optimization探索过程中使用白盒模型作为指导。使用行业标准基准应用程序对Apache Spark进行的评估表明,GBO始终在应用程序工作负载中提供性能加速,节省的幅度接近2倍。
There is a lot of interest today in building autonomous (or, self-driving) data processing systems. An emerging school of thought is to leverage the "black box" algorithm of Bayesian Optimization for problems of this flavor both due to its wider applicability and theoretical guarantees on the quality of results produced. The black-box approach, however, could be time and labor-intensive; or otherwise get stuck in a local minima. We study an important problem of auto-tuning the memory allocation for applications running on modern distributed data processing systems. A simple "white-box" model is developed which can quickly separate good configurations from bad ones. To combine the benefits of the two approaches to tuning, we build a framework called Guided Bayesian Optimization (GBO) that uses the white-box model as a guide during the Bayesian Optimization exploration process. An evaluation carried out on Apache Spark using industry-standard benchmark applications shows that GBO consistently provides performance speedups across the application workload with the magnitude of savings being close to 2x.