AT-GIS: Highly Parallel Spatial Query Processing with Associative Transducers

AT-GIS: Highly Parallel Spatial Query Processing with Associative Transducers
复制标题

AT-GIS:使用关联传感器的高度并行空间查询处理

DOI:
--
复制
发表时间:
2016
期刊:
SIGMOD Conference
影响因子:
--
通讯作者:
P. Pietzuch
P. Pietzuch
中科院分区:
--
文献类型:
--
作者:
Peter Ogden;David B. Thomas;P. Pietzuch

文献摘要

被引文献

相似文献

城市规划、交通和环境科学等领域的用户都希望对不断更新的空间数据集执行分析查询。当前的大规模空间查询处理解决方案要么依赖于对RDBMS的扩展,这在数据更改时需要昂贵的加载和索引阶段,要么依赖于分布式映射/缩减框架,在资源匮乏的计算集群上运行。这两种解决方案都要解决解析复杂的、层次化的空间数据格式的顺序瓶颈问题,这通常会占用查询执行时间。我们的目标是充分利用现代多核CPU为解析和查询执行提供的并行性,从而提供具有单机资源的集群性能。我们描述了AT-GIS,一个高度并行的空间查询处理系统,线性扩展到大量的CPU核心。AT-GIS集成了解析和查询的空间数据使用一种新的计算抽象称为关联传感器(AT)。AT可以形成用于计算的单个数据并行流水线,而不需要将空间输入数据分成逻辑上独立的块。使用AT,AT-GIS可以并行地对多种格式的原始输入数据执行空间查询操作,而无需任何预处理。在单个64核机器上,AT-GIT提供了8节点Hadoop集群的3倍性能,其中192个核心用于包含查询,10倍用于聚合查询。
Users in many domains, including urban planning, transportation, and environmental science want to execute analytical queries over continuously updated spatial datasets. Current solutions for large-scale spatial query processing either rely on extensions to RDBMS, which entails expensive loading and indexing phases when the data changes, or distributed map/reduce frameworks, running on resource-hungry compute clusters. Both solutions struggle with the sequential bottleneck of parsing complex, hierarchical spatial data formats, which frequently dominates query execution time. Our goal is to fully exploit the parallelism offered by modern multi-core CPUs for parsing and query execution, thus providing the performance of a cluster with the resources of a single machine. We describe AT-GIS, a highly-parallel spatial query processing system that scales linearly to a large number of CPU cores. AT-GIS integrates the parsing and querying of spatial data using a new computational abstraction called associative transducers (ATs). ATs can form a single data-parallel pipeline for computation without requiring the spatial input data to be split into logically independent blocks. Using ATs, AT-GIS can execute, in parallel, spatial query operators on the raw input data in multiple formats, without any pre-processing. On a single 64-core machine, AT-GIT provides 3x the performance of an 8-node Hadoop cluster with 192 cores for containment queries, and 10x for aggregation queries.