Profiling and Predicting Application Performance on the Cloud

Profiling and Predicting Application Performance on the Cloud
复制标题

DOI:
10.1109/ucc.2018.00011
复制
发表时间:
2018-12
期刊:
2018 IEEE/ACM 11th International Conference on Utility and Cloud Computing (UCC)
影响因子:
--
通讯作者:
Matt Baughman;Ryan Chard;Logan T. Ward;Jason Pitt;K. Chard;Ian T Foster
Matt Baughman;Ryan Chard;Logan T. Ward;Jason Pitt;K. Chard;Ian T Foster
中科院分区:
其他
文献类型:
--
作者:
Matt Baughman;Ryan Chard;Logan T. Ward;Jason Pitt;K. Chard;Ian T Foster

文献摘要

相似文献

云提供商不断扩展其可租赁资源集合并使其多样化,以满足日益广泛的应用程序的需求。虽然这种灵活性是云的一个关键优势,但它也造成了一个复杂的环境,用户在给定的应用程序中面临着多种资源选择。次优选择会降低性能并增加成本。由于资源池的快速发展,用户无法单独选择实例类型;相反,需要自动化方法来简化和指导资源配置。在这里,我们提出了一种自动预测任意云实例上的应用程序性能的方法。我们结合离线和在线分析方法,使用从非云环境收集的历史数据和在云环境上运行的目标分析来创建一个复合应用程序模型,该模型可以预测给定输入数据大小的给定云实例类型上的运行时间。我们证明了生产生物信息学工作流程中使用的九个应用程序的平均错误率为 17.2%。最后,我们评估了一种实验设计方法,以探索分析成本和模型准确性之间的权衡。使用这种方法,在没有先验知识的情况下,我们证明,使用 4 个选择性实验,我们可以实现使用所有实例类型训练的模型的 30% 以内的性能。
Cloud providers continue to expand and diversify their collection of leasable resources to meet the needs of an increasingly wide range of applications. While this flexibility is a key benefit of the cloud, it also creates a complex landscape in which users are faced with many resource choices for a given application. Suboptimal selections can both degrade performance and increase costs. Given the rapidly evolving pool of resources, it is infeasible for users alone to select instance types; instead, automated methods are needed to simplify and guide resource provisioning. Here we present a method for the automatic prediction of application performance on arbitrary cloud instances. We combine offline and online profiling approaches, using historical data gathered from non-cloud environments and targeted profiling runs on cloud environments to create a composite application model that can predict run times on a given cloud instance type for a given input data size. We demonstrate average error of 17.2% across nine applications used in production bioinformatics workflows. Finally, we evaluate an experiment design approach to explore the trade-off between the cost of profiling and the accuracy of our models. Using this approach, with no prior knowledge, we show that using 4 selectively chosen experiments we can achieve performance within 30% of a model trained using all instance types.