On Spatial Joins in MapReduce

On Spatial Joins in MapReduce
复制标题

MapReduce 中的空间连接

DOI:
--
复制
发表时间:
2017
期刊:
SIGSPATIAL/GIS
影响因子:
--
通讯作者:
M. Mokbel
M. Mokbel
中科院分区:
--
文献类型:
--
作者:
Ibrahim Sabek;M. Mokbel

文献摘要

被引文献

相似文献

本文为基于mapreduce的空间连接算法的成熟查询优化器提供了第一次尝试。优化器开发了自己的分类法,该分类法几乎涵盖了对任意两个输入数据集进行空间连接的所有可能方法。优化器有两种形式;基于成本和基于规则。给定两个输入数据集,基于成本的查询优化器评估已开发分类法中所有可能选项的成本,并选择成本最低的选项。基于规则的查询优化器将基于成本优化器开发的成本模型抽象为一组简单的易于检查的启发式规则。然后,应用它的规则来选择成本最低的选项。这两个查询优化器都在一个广泛使用的基于开源mapreduce的大空间数据系统中进行了部署和实验评估。详尽的实验表明,这两个查询优化器在空间连接任何两个数据集(每个数据集最多500GB)时总是能够成功地做出正确的决策。
This paper provides the first attempt for a full-fledged query optimizer for MapReduce-based spatial join algorithms. The optimizer develops its own taxonomy that covers almost all possible ways of doing a spatial join for any two input datasets. The optimizer comes in two flavors; cost-based and rule-based. Given two input data sets, the cost-based query optimizer evaluates the costs of all possible options in the developed taxonomy, and selects the one with the lowest cost. The rule-based query optimizer abstracts the developed cost models of the cost-based optimizer into a set of simple easy-to-check heuristic rules. Then, it applies its rules to select the lowest cost option. Both query optimizers are deployed and experimentally evaluated inside a widely used open-source MapReduce-based big spatial data system. Exhaustive experiments show that both query optimizers are always successful in taking the right decision for spatially joining any two datasets of up to 500GB each.