SWARM: Adaptive Load Balancing in Distributed Streaming Systems for Big Spatial Data

SWARM: Adaptive Load Balancing in Distributed Streaming Systems for Big Spatial Data
复制标题

DOI:
10.1145/3460013
复制
发表时间:
2020-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Anas Daghistani;W. Aref;A. Ghafoor;Ahmed R. Mahmood
Anas Daghistani;W. Aref;A. Ghafoor;Ahmed R. Mahmood
中科院分区:
其他
文献类型:
--
作者:
Anas Daghistani;W. Aref;A. Ghafoor;Ahmed R. Mahmood

文献摘要

相似文献

支持GPS的设备的扩散导致了许多基于位置的服务的开发导致分布式空间流系统的开发。相比之下,现有的空间分区来分配工作负载。为了应对空间数据和查询的分布的变化。性能瓶颈会被检测到多个查询执行和数据持久性模型。使用真实和合成数据集进行了广泛的实验评估,表明,基于观察者确定的静态网格分区的吞吐量平均而言,吞吐量有限,而数据和查询工作量有限。与其他技术相比。
The proliferation of GPS-enabled devices has led to the development of numerous location-based services. These services need to process massive amounts of streamed spatial data in real-time. The current scale of spatial data cannot be handled using centralized systems. This has led to the development of distributed spatial streaming systems. Existing systems are using static spatial partitioning to distribute the workload. In contrast, the real-time streamed spatial data follows non-uniform spatial distributions that are continuously changing over time. Distributed spatial streaming systems need to react to the changes in the distribution of spatial data and queries. This article introduces SWARM, a lightweight adaptivity protocol that continuously monitors the data and query workloads across the distributed processes of the spatial data streaming system and redistributes and rebalances the workloads as soon as performance bottlenecks get detected. SWARM is able to handle multiple query-execution and data-persistence models. A distributed streaming system can directly use SWARM to adaptively rebalance the system’s workload among its machines with minimal changes to the original code of the underlying spatial application. Extensive experimental evaluation using real and synthetic datasets illustrate that, on average, SWARM achieves 2 improvement in throughput over a static grid partitioning that is determined based on observing a limited history of the data and query workloads. Moreover, SWARM reduces execution latency on average 4 compared with the other technique.