Adaptive Fragment-Based Parallel State Recovery for Stream Processing Systems

Adaptive Fragment-Based Parallel State Recovery for Stream Processing Systems
复制标题

DOI:
10.1109/tpds.2023.3251997
复制
发表时间:
2023-08
影响因子:
5.3
通讯作者:
Hailu Xu;Pinchao Liu;Sarker Tanzir Ahmed;Dilma Da Silva;Liting Hu
Hailu Xu;Pinchao Liu;Sarker Tanzir Ahmed;Dilma Da Silva;Liting Hu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hailu Xu;Pinchao Liu;Sarker Tanzir Ahmed;Dilma Da Silva;Liting Hu

文献摘要

相似文献

如今,大规模的云组织正在部署数据中心,并且在全球范围内提供了对服务的低延迟访问权限。处理,留下另一个紧急的趋势流处理,探索较少。重要的是,当现有研究发生故障和失败时,成功地恢复了大型分布状态提供状态恢复主要通过三种方法提供:复制恢复,检查点恢复和基于Dstream的谱系恢复,它们要么很慢,要么资源昂贵或无法处理多个故障;(3)它们不适合异质性硬件设置。将基于哈希表的分布式对等覆盖层中,并将每个节点的局部状态分为许多片段。不同的片段可以并行重建失败状态。与Apache Storm相比,A-FP4使用现实世界数据集的大规模实验降低了31.8%至50.5%。演示A-FP4的有吸引力的可伸缩性和适应性属性。
Today, large-scale cloud organizations are deploying datacenters and “edge” clusters globally to provide low-latency access to services. Running stream applications across geo-distributed sites are emerging as a daily requirement. However, existing efforts have dominantly centered around stateless stream processing, leaving another urgent trend-stateful stream processing-much less explored. A driving need is to store and update states during processing, and most importantly, successfully recover large distributed states when faults and failures happen. Existing studies exhibit major limitations including: (1) they mostly inherit MapReduce's “single master/many workers” architecture, where the central master can easily become ascalability bottleneck; (2) they offer state recovery mainly through three approaches: replication recovery, checkpointing recovery, and DStream-based lineage recovery, which are either slow, resource-expensive or failing to handle multiple failures; and (3) they are not adaptive to heterogeneous hardware settings. We present A-FP4S, a novel adaptive fragments-based parallel state recovery mechanism for stream processing systems. A-FP4S organizes stream operators into a distributed hash table based peer-to-peer overlay and divides each node's local state into many fragments. These fragments are periodically stored in node's multiple neighbors, ensuring different sets of available fragments can reconstruct failed states in parallel. This mechanism is extremely scalable to the lost state, significantly reduces failure recovery time, and can tolerate multiple node failures. A-FP4S is adaptive to heterogeneous hardware settings by automatic parameter tuning over phases. Compared to Apache Storm, A-FP4S achieves 31.8% to 50.5% reduction in recovery latency. Large-scale experiments using real-world datasets demonstrate A-FP4S's attractive scalability and adaptivity properties.