Network Cost-Aware Geo-Distributed Data Analytics System

Network Cost-Aware Geo-Distributed Data Analytics System
复制标题

DOI:
10.1109/tpds.2021.3108893
复制
发表时间:
2021-08
影响因子:
5.3
通讯作者:
Kwangsung Oh;Minmin Zhang;A. Chandra;J. Weissman
Kwangsung Oh;Minmin Zhang;A. Chandra;J. Weissman
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kwangsung Oh;Minmin Zhang;A. Chandra;J. Weissman

文献摘要

被引文献

相似文献

许多地理分布式数据分析(GDA)系统都集中在网络性能瓶颈:数据中心间的网络带宽,以提高性能。不幸的是,这些系统可能会遇到成本瓶颈(${\$}$$),因为它们没有考虑数据传输成本(${\$}$$),这是多云环境中最昂贵的异构资源之一。在这篇文章中,我们提出了泡菜,网络成本意识的GDA系统,以满足成本性能权衡利用数据传输成本的异质性,以避免成本瓶颈。Kimchi确定用于调度任务的具有成本意识的任务布局决策,给定输入包括数据传输成本、网络带宽、输入数据大小和位置以及期望的成本性能权衡偏好。此外,Kimchi还注意到动态情况下的数据传输成本。Kimchi已经应用于两种常见的GDA MapReduce模型:同步屏障和基于异步推送的shuffle。Kimchi原型已经在Spark上实现,实验表明,与其他基线方法集中式,vanilla Spark和带宽感知(例如,铱)。更重要的是,Kimchi允许应用程序在多云环境中探索更丰富的性价比权衡空间。
Many geo-distributed data analytics (GDA) systems have focused on the network performance-bottleneck: inter-data center network bandwidth to improve performance. Unfortunately, these systems may encounter a cost-bottleneck (${\$}$$) because they have not considered data transfer cost (${\$}$$), one of the most expensive and heterogeneous resources in a multi-cloud environment. In this article, we present Kimchi, a network cost-aware GDA system to meet the cost-performance tradeoff by exploiting data transfer cost heterogeneity to avoid the cost-bottleneck. Kimchi determines cost-aware task placement decisions for scheduling tasks given inputs including data transfer cost, network bandwidth, input data size and locations, and desired cost-performance tradeoff preference. In addition, Kimchi is also mindful of data transfer cost in the presence of dynamics. Kimchi has been applied to two common GDA MapReduce models: synchronous barrier and asynchronous push-based shuffle. A Kimchi prototype has been implemented on Spark, and experiments show that it reduces cost by 5% $\scriptstyle \sim$∼ 24% without impacting performance and reduces query execution time by 45% $\scriptstyle \sim$∼ 70% without impacting cost compared to other baseline approaches centralized, vanilla Spark, and bandwidth-aware (e.g., Iridium). More importantly, Kimchi allows applications to explore a much richer cost-performance tradeoff space in a multi-cloud environment.