Plexus: Optimizing Join Approximation for Geo-Distributed Data Analytics

Plexus: Optimizing Join Approximation for Geo-Distributed Data Analytics
复制标题

Plexus:优化地理分布式数据分析的连接近似

DOI:
10.1145/3620678.3624643
复制
发表时间:
2023
期刊:
IC2E
影响因子:
--
通讯作者:
Chandra, Abhishek
Chandra, Abhishek
中科院分区:
--
文献类型:
--
作者:
Wolfrath, Joel;Chandra, Abhishek

文献摘要

参考文献

被引文献

相似文献

现代应用程序越来越多地跨地理分布式数据中心或边缘集群而不是单个云生成和持久化数据。这种范例给传统的查询执行带来了挑战,因为在广域网链路上传输数据时会增加延迟。连接查询受到的影响尤其严重,因为它们的输出规模很大,而且必须在网络上传输大量数据。连接采样——从连接结果中计算一个统一的样本——是一种减少资源需求的有用技术。然而,将其应用于地理分布设置是具有挑战性的,因为从每个位置获取独立的样本并连接样本并不能从连接结果产生统一且独立的元组。为了解决这些挑战,我们首先将现有的连接采样算法推广到地理分布设置。然后我们介绍了我们的系统Plexus,它引入了三个额外的优化来进一步减少网络开销并处理网络和数据异构:(i)权重近似,(ii)异构感知和(iii)样本预取。我们在一个部署在多个AWS区域的地理分布式系统上对Plexus进行了评估,该系统基于Apache Spark实现。通过使用三个真实的数据集,我们展示了Plexus可以在广泛的连接查询类别上比默认的Spark连接实现减少高达80%的查询延迟,而不会显著影响样本一致性。
Modern applications are increasingly generating and persisting data across geo-distributed data centers or edge clusters rather than a single cloud. This paradigm introduces challenges for traditional query execution due to increased latency when transferring data over wide-area network links. Join queries in particular are heavily affected, due to their large output size and amount of data that must be shuffled over the network. Join sampling---computing a uniform sample from the join results---is a useful technique for reducing resource requirements. However, applying it to a geo-distributed setting is challenging, since acquiring independent samples from each location and joining on the samples does not produce uniform and independent tuples from the join result. To address these challenges, we first generalize an existing join sampling algorithm to the geo-distributed setting. We then present our system, Plexus, which introduces three additional optimizations to further reduce the network overhead and handle network and data heterogeneity: (i) weight approximation, (ii) heterogeneity awareness and (iii) sample prefetching. We evaluate Plexus on a geo-distributed system deployed across multiple AWS regions, with an implementation based on Apache Spark. Using three real-world datasets, we show that Plexus can reduce query latency by up to 80% over the default Spark join implementation on a wide class of join queries without substantially impacting sample uniformity.
AggFirstJoin:使用基于聚合的转换优化地理分布式连接
DOI: 10.1109/ccgrid57682.2023.00046
发表时间: 2023
期刊: Cloud and Internet Computing (CCGrid
影响因子: --
作者:
Kumar, Dhruv;Ahmad, Sohaib;Chandra, Abhishek;Sitaraman, Ramesh K.
通讯作者: Sitaraman, Ramesh K.
DOI: 10.1109/ic2e55432.2022.00013
发表时间: 2022-08
期刊: 2022 IEEE International Conference on Cloud Engineering (IC2E)
影响因子: --
作者:
Joel Wolfrath;A. Chandra
通讯作者: Joel Wolfrath;A. Chandra
DOI: 10.1145/3267809.3267834
发表时间: 2018-10
期刊: Proceedings of the ACM Symposium on Cloud Computing
影响因子: --
作者:
D. Quoc;Istemi Ekin Akkus;Pramod Bhatotia;Spyros Blanas;Ruichuan Chen;C. Fetzer;T. Strufe
通讯作者: D. Quoc;Istemi Ekin Akkus;Pramod Bhatotia;Spyros Blanas;Ruichuan Chen;C. Fetzer;T. Strufe
针对地理分布式数据进行广域网感知连接采样
DOI: 10.1145/3517206.3526268
发表时间: 2022
期刊: Analytics and Networking
影响因子: --
作者:
Kumar, Dhruv;Wolfrath, Joel;Chandra, Abhishek;Sitaraman, Ramesh K.
通讯作者: Sitaraman, Ramesh K.
数据仓库的高效连接概要维护
DOI: --
发表时间: 2020
期刊: SIGMOD Conference
影响因子: --
作者:
Zhuoyue Zhao;Feifei Li;Yuxi Liu
通讯作者: Yuxi Liu