Smartpick: Workload Prediction for Serverless-enabled Scalable Data Analytics Systems

Smartpick: Workload Prediction for Serverless-enabled Scalable Data Analytics Systems
复制标题

DOI:
10.1145/3590140.3592850
复制
发表时间:
2023-07
期刊:
Proceedings of the 24th International Middleware Conference
影响因子:
--
通讯作者:
Anshuman Mohapatra;Kwangsung Oh
Anshuman Mohapatra;Kwangsung Oh
中科院分区:
其他
文献类型:
--
作者:
Anshuman Mohapatra;Kwangsung Oh

文献摘要

相似文献

许多数据分析系统已经采用了一种新出现的计算资源,即无服务器(SL),以及时和经济高效的方式处理数据分析查询,即,无服务器数据分析虽然由于SL的灵活性和可扩展性,这些系统可以快速开始处理查询,但由于SL的性能比传统计算资源更差,成本更昂贵,因此它们可能会遇到基于工作负载的性能和成本瓶颈,例如,虚拟机(VM)。在本文中,我们介绍了Smartpick,这是一个支持SL的可扩展数据分析系统,它将SL和VM结合在一起,以实现综合效益,即,SL带来的灵活性以及VM带来的更高性能和更低成本。Smartpick使用机器学习预测方案,基于决策树的随机森林和贝叶斯优化器,来确定SL和VM配置,即,有多少SL和VM实例用于查询,以满足成本性能目标。Smartpick为应用程序提供了一个旋钮,允许它们探索通过同时利用SL和VM打开的更丰富的性价比权衡空间。为了最大限度地发挥SL的优势,Smartpick支持一种简单但强大的机制,称为中继实例。Smartpick还支持事件驱动的预测模型再训练,以处理工作负载动态。Smartpick原型在Spark上实现,并部署在实时测试平台、Amazon AWS和Google Cloud Platform上。评估结果表明,97.05%和83.49%的预测精度分别高达50%的成本降低,而不是基线。结果还证实,Smartpick允许数据分析应用程序有效地导航更丰富的性价比权衡空间,并有效地自动处理工作负载动态。
Many data analytic systems have adopted a newly emerging compute resource, serverless (SL), to handle data analytics queries in a timely and cost-efficient manner, i.e., serverless data analytics. While these systems can start processing queries quickly thanks to the agility and scalability of SL, they may encounter performance-and cost-bottlenecks based on workloads due to SL's worse performance and more expensive cost than traditional compute resources, e.g., virtual machine (VM). In this paper, we introduce Smartpick, a SL-enabled scalable data analytics system that exploits SL and VM together to realize composite benefits, i.e., agility from SL and better performance with reduced cost from VM. Smartpick uses a machine learning prediction scheme, decision-tree based Random Forest with Bayesian Optimizer, to determine SL and VM configurations, i.e., how many SL and VM instances for queries, that meet cost-performance goals. Smartpick offers a knob for applications to allow them to explore a richer cost-performance tradeoff space opened by exploiting SL and VM together. To maximize the benefits of SL, Smartpick supports a simple but strong mechanism, called relay-instances. Smartpick also supports event-driven prediction model retraining to deal with workload dynamics. A Smartpick prototype was implemented on Spark and deployed on live testbeds, Amazon AWS and Google Cloud Platform. Evaluation results indicate 97.05% and 83.49% prediction accuracies respectively with up to 50% cost reduction as opposed to the baselines. The results also confirm that Smartpick allows data analytics applications to navigate the richer cost-performance tradeoff space efficiently and to handle workload dynamics effectively and automatically.