AggFirstJoin: Optimizing Geo-Distributed Joins using Aggregation-Based Transformations

AggFirstJoin: Optimizing Geo-Distributed Joins using Aggregation-Based Transformations
复制标题

AggFirstJoin:使用基于聚合的转换优化地理分布式连接

DOI:
10.1109/ccgrid57682.2023.00046
复制
发表时间:
2023
期刊:
Cloud and Internet Computing (CCGrid
影响因子:
--
通讯作者:
Sitaraman, Ramesh K.
Sitaraman, Ramesh K.
中科院分区:
--
文献类型:
--
作者:
Kumar, Dhruv;Ahmad, Sohaib;Chandra, Abhishek;Sitaraman, Ramesh K.

文献摘要

参考文献

被引文献

相似文献

地理分布分析(GDA)涉及处理跨地理分布的站点存储的数据。此类分析涉及广域网(WAN)链路上的数据传输。广域网链路具有高度受限和异构性,这使得广域网中的数据传输速度慢且成本高。为了解决这个问题,最近的方法已经提出了广域网感知的调度和地理分布式分析任务的放置。然而,地理分布环境中的计算联接仍然是一个具有挑战性的问题。在这项工作中,我们提出了AggFirstJoin,一种使用理论上合理的查询转换技术来最小化地理分布连接的代价的方法。我们的优化方法综合考虑了连接和聚合操作,这些操作通常是同一查询的一部分,并在连接之前推送(转换后的)聚合,以产生与原始查询相同的结果。我们使用支持广域网的任务放置和Bloom过滤方法来增强查询转换技术,以分别进一步减少查询执行时间和广域网使用量。我们在一个流行的大数据分析引擎--ApacheSpark之上实现了我们提出的技术。我们使用合成、TPC-H和AMPLab大数据基准数据集在AWS上的真实地理分布测试床和模拟测试床上对我们提出的技术进行了广泛的评估。我们的评估表明,与最先进的GDA技术相比,我们提出的技术将查询执行时间减少了300倍,广域网使用量减少了200倍。
Geo-distributed analytics (GDA) involves processing of data stored across geographically distributed sites. Such analytics involves data transfer over the wide area network (WAN) links. WAN links are highly constrained and heterogeneous in nature, making the data transfer over the WAN slow and costly. To tackle this issue, recent approaches have proposed WAN-aware scheduling and placement of geo-distributed analytics tasks. However, computing joins in a geo-distributed setting remains a challenging problem. In this work, we propose AggFirstJoin, an approach to minimize the cost of geo-distributed joins using a theoretically sound query transformation technique. Our optimization approach takes a combined view of the join and aggregation operations which are often part of the same query and pushes (a transformed) aggregation before join in a manner to produce the same results as the original query. We augment our query transformation technique with a WAN-aware task placement and a Bloom filtering approach to further reduce query execution time and WAN usage respectively. We implement our proposed technique on top of Apache Spark, a popular engine for big data analytics. We extensively evaluate our proposed technique using synthetic, TPC-H and Amplab Big Data benchmark datasets on a real geo-distributed testbed on AWS as well as an emulated testbed. Our evaluations show our proposed technique achieves up to 300x reduction in query execution time and 200x reduction in WAN usage as compared to state-of-the-art GDA techniques.
用于决策支持查询的位向量感知查询优化
DOI: --
发表时间: 2020
期刊: SIGMOD Conference
影响因子: --
作者:
B. Ding;S. Chaudhuri;Vivek R. Narasayya
通讯作者: Vivek R. Narasayya
DOI: 10.1145/3423211.3425668
发表时间: 2020-12
期刊: Proceedings of the 21st International Middleware Conference
影响因子: --
作者:
A. Jonathan;A. Chandra;J. Weissman
通讯作者: A. Jonathan;A. Chandra;J. Weissman
DOI: --
发表时间: 1994-09
期刊: --
影响因子: --
作者:
S. Chaudhuri;Kyuseok Shim
通讯作者: S. Chaudhuri;Kyuseok Shim
DOI: 10.1145/3309697.3331491
发表时间: 2019-06
期刊: Abstracts of the 2019 SIGMETRICS/Performance Joint International Conference on Measurement and Modeling of Computer Systems
影响因子: --
作者:
Dhruv Kumar;Jian Li;A. Chandra;R. Sitaraman
通讯作者: Dhruv Kumar;Jian Li;A. Chandra;R. Sitaraman
DOI: 10.1145/3267809.3267834
发表时间: 2018-10
期刊: Proceedings of the ACM Symposium on Cloud Computing
影响因子: --
作者:
D. Quoc;Istemi Ekin Akkus;Pramod Bhatotia;Spyros Blanas;Ruichuan Chen;C. Fetzer;T. Strufe
通讯作者: D. Quoc;Istemi Ekin Akkus;Pramod Bhatotia;Spyros Blanas;Ruichuan Chen;C. Fetzer;T. Strufe