State Management in Apache Flink®: Consistent Stateful Distributed Stream Processing

State Management in Apache Flink®: Consistent Stateful Distributed Stream Processing
复制标题

DOI:
10.14778/3137765.3137777
复制
发表时间:
2017-08
期刊:
Proc. VLDB Endow.
影响因子:
--
通讯作者:
Paris Carbone;Stephan Ewen;Gyula Fóra;Seif Haridi;Stefan Richter;K. Tzoumas
Paris Carbone;Stephan Ewen;Gyula Fóra;Seif Haridi;Stefan Richter;K. Tzoumas
中科院分区:
其他
文献类型:
--
作者:
Paris Carbone;Stephan Ewen;Gyula Fóra;Seif Haridi;Stefan Richter;K. Tzoumas

文献摘要

被引文献

相似文献

流处理器正在工业中作为一种设备出现,它驱动处理持久应用逻辑核心的分析服务,但也是关键任务服务。因此,除了可扩展性和低延迟之外,日益增长的系统需求是对应用程序状态的一流支持以及强大的一致性保证,以及对群集重新配置、软件补丁和部分故障的适应性。虽然以前的系统研究已经解决了其中一些具体问题,但实际的挑战在于如何以透明、非侵入性的方式实现这种保证,使用户摆脱不必要的限制。这样的需求是一个开源的、可伸缩的流处理器ApacheFlink中状态管理的主要设计原则。我们介绍了Flink的核心流水线动态机制,它保证逐步创建轻量级的、一致的、分布式的应用程序状态快照,而不会影响持续执行。一致的快照通过粗粒度回滚恢复,覆盖系统重新配置、容错和版本管理的所有需求。应用程序状态向系统显式声明,从而允许高效分区和对持久存储的透明提交。我们进一步介绍了Flink的高可用性、外部状态查询和输出提交的后端实现和机制。最后,我们用指标和大型部署洞察展示了这些机制在实践中的行为,展示了我们方法的低性能权衡,以及在连续但可持续的系统部署中利用异步性的一般好处。
Stream processors are emerging in industry as an apparatus that drives analytical but also mission critical services handling the core of persistent application logic. Thus, apart from scalability and low-latency, a rising system need is first-class support for application state together with strong consistency guarantees, and adaptivity to cluster reconfigurations, software patches and partial failures. Although prior systems research has addressed some of these specific problems, the practical challenge lies on how such guarantees can be materialized in a transparent, non-intrusive manner that relieves the user from unnecessary constraints. Such needs served as the main design principles of state management in Apache Flink, an open source, scalable stream processor. We present Flink's core pipelined, in-flight mechanism which guarantees the creation of lightweight, consistent, distributed snapshots of application state, progressively, without impacting continuous execution. Consistent snapshots cover all needs for system reconfiguration, fault tolerance and version management through coarse grained rollback recovery. Application state is declared explicitly to the system, allowing efficient partitioning and transparent commits to persistent storage. We further present Flink's backend implementations and mechanisms for high availability, external state queries and output commit. Finally, we demonstrate how these mechanisms behave in practice with metrics and large-deployment insights exhibiting the low performance trade-offs of our approach and the general benefits of exploiting asynchrony in continuous, yet sustainable system deployments.