Is Big Data Performance Reproducible in Modern Cloud Networks?

Is Big Data Performance Reproducible in Modern Cloud Networks?
复制标题

DOI:
--
复制
发表时间:
2019-12
期刊:
--
影响因子:
--
通讯作者:
Alexandru Uta;Alexandru Custura;Dmitry Duplyakin;I. Jimenez;Jan S. Rellermeyer;C. Maltzahn;R. Ricci;A. Iosup
Alexandru Uta;Alexandru Custura;Dmitry Duplyakin;I. Jimenez;Jan S. Rellermeyer;C. Maltzahn;R. Ricci;A. Iosup
中科院分区:
其他
文献类型:
--
作者:
Alexandru Uta;Alexandru Custura;Dmitry Duplyakin;I. Jimenez;Jan S. Rellermeyer;C. Maltzahn;R. Ricci;A. Iosup

文献摘要

被引文献

相似文献

十多年来,性能可变性一直被云从业者和性能工程师视为一个问题。然而,我们对顶级系统会议的调查显示,研究界在云中进行实验时经常忽视可变性。以网络为重点,我们通过从主流商业云和私有研究云收集跟踪数据来评估可变性对基于云的大数据工作负载的影响。我们的数据收集包含在传输超过9拍字节数据时收集的数百万个数据点。我们描述了数据中存在的网络可变性,并表明,尽管商业云提供商实施了服务质量保障机制,但可变性仍然存在,而且此类机制和服务提供商的策略甚至会加剧这种可变性。我们展示了大数据工作负载如何出现显著的减速,并且缺乏可预测性和可复制性,即使使用了最先进的实验技术也是如此。我们为从业者提供了减少大数据性能波动性的指南,使实验更具可重复性。
Performance variability has been acknowledged as a problem for over a decade by cloud practitioners and performance engineers. Yet, our survey of top systems conferences reveals that the research community regularly disregards variability when running experiments in the cloud. Focusing on networks, we assess the impact of variability on cloud-based big-data workloads by gathering traces from mainstream commercial clouds and private research clouds. Our data collection consists of millions of datapoints gathered while transferring over 9 petabytes of data. We characterize the network variability present in our data and show that, even though commercial cloud providers implement mechanisms for quality-of-service enforcement, variability still occurs, and is even exacerbated by such mechanisms and service provider policies. We show how big-data workloads suffer from significant slowdowns and lack predictability and replicability, even when state-of-the-art experimentation techniques are used. We provide guidelines for practitioners to reduce the volatility of big data performance, making experiments more repeatable.