GeoFlink: A Distributed and Scalable Framework for the Real-time Processing of Spatial Streams

GeoFlink: A Distributed and Scalable Framework for the Real-time Processing of Spatial Streams
复制标题

DOI:
10.1145/3340531.3412761
复制
发表时间:
2020-04
期刊:
Proceedings of the 29th ACM International Conference on Information & Knowledge Management
影响因子:
--
通讯作者:
Salman Ahmed Shaikh;Komal Mariam;H. Kitagawa;Kyoung-Sook Kim
Salman Ahmed Shaikh;Komal Mariam;H. Kitagawa;Kyoung-Sook Kim
中科院分区:
其他
文献类型:
--
作者:
Salman Ahmed Shaikh;Komal Mariam;H. Kitagawa;Kyoung-Sook Kim

文献摘要

相似文献

Apache Flink是一个开源系统,用于批处理和流数据的可扩展处理。Flink本身不支持空间数据流的高效处理,而这是许多处理空间数据的应用程序所需要的。除了Flink,其他可扩展的空间数据处理平台包括GeoSpark、spatial Hadoop等都不支持流工作负载,只能处理静态/批处理工作负载。为了填补这一空白,我们提出了GeoFlink,它扩展了Apache Flink,以支持空间数据类型、索引和对空间数据流的连续查询。为了实现空间连续查询的高效处理和Flink集群节点间数据的有效分布,引入了基于网格的索引。GeoFlink目前支持点数据类型的空间范围、空间kNN和空间连接查询。对真实空间数据流的实验研究表明,GeoFlink的查询吞吐量明显高于普通Flink处理。
Apache Flink is an open-source system for scalable processing of batch and streaming data. Flink does not natively support efficient processing of spatial data streams, which is a requirement of many applications dealing with spatial data. Besides Flink, other scalable spatial data processing platforms including GeoSpark, Spatial Hadoop, etc. do not support streaming workloads and can only handle static/batch workloads. To fill this gap, we present GeoFlink, which extends Apache Flink to support spatial data types, indexes and continuous queries over spatial data streams. To enable efficient processing of spatial continuous queries and for the effective data distribution across Flink cluster nodes, a gird-based index is introduced. GeoFlink currently supports spatial range, spatial kNN and spatial join queries on point data type. An experimental study on real spatial data streams shows that GeoFlink achieves significantly higher query throughput than ordinary Flink processing.