ConEx: Efficient Exploration of Big-Data System Configurations for Better Performance

ConEx: Efficient Exploration of Big-Data System Configurations for Better Performance
复制标题

DOI:
10.1109/tse.2020.3007560
复制
发表时间:
2022-03-01
影响因子:
7.4
通讯作者:
Ray, Baishakhi
Ray, Baishakhi
中科院分区:
计算机科学1区
文献类型:
--
作者:
Krishna, Rahul;Tang, Chong;Ray, Baishakhi

文献摘要

被引文献

相似文献

配置空间的复杂性使得大数据软件系统难以很好地配置。考虑 Hadoop,它有超过 900 个参数,开发人员通常只使用 Hadoop 发行版提供的默认配置。绩效损失的机会成本是巨大的。由于收集训练数据的成本很高,流行的基于学习的自动调整软件方法对于大数据系统来说不能很好地扩展。我们提出了一种基于进化马尔可夫链蒙特卡罗 (EMCMC) 采样和成本降低技术相结合的新方法,为大数据系统找到性能更好的配置。为了降低成本,我们开发并实验测试和验证了两种方法:使用扩大的大数据作业作为较大作业的目标函数的代理,并使用动态作业相似性度量来推断针对一种大数据问题获得的结果将很好地适用于类似的问题。我们的实验结果表明,我们的方法有望显着提高大数据系统的性能,并且优于基于随机采样、基本遗传算法(GA)和预测模型学习的竞争方法。我们的实验结果支持这样的结论:我们的方法有力地证明了显着且节俭地提高大数据系统性能的潜力。
Configuration space complexity makes the big-data software systems hard to configure well. Consider Hadoop, with over nine hundred parameters, developers often just use the default configurations provided with Hadoop distributions. The opportunity costs in lost performance are significant. Popular learning-based approaches to auto-tune software does not scale well for big-data systems because of the high cost of collecting training data. We present a new method based on a combination of Evolutionary Markov Chain Monte Carlo (EMCMC) sampling and cost reduction techniques to find better-performing configurations for big data systems. For cost reduction, we developed and experimentally tested and validated two approaches: using scaled-up big data jobs as proxies for the objective function for larger jobs and using a dynamic job similarity measure to infer that results obtained for one kind of big data problem will work well for similar problems. Our experimental results suggest that our approach promises to improve the performance of big data systems significantly and that it outperforms competing approaches based on random sampling, basic genetic algorithms (GA), and predictive model learning. Our experimental results support the conclusion that our approach strongly demonstrates the potential to improve the performance of big data systems significantly and frugally.