Highly available, fault-tolerant, parallel dataflows

Highly available, fault-tolerant, parallel dataflows
复制标题

DOI:
10.1145/1007568.1007662
复制
发表时间:
2004-06
期刊:
--
影响因子:
--
通讯作者:
Mehul A. Shah;J. Hellerstein;E. Brewer
Mehul A. Shah;J. Hellerstein;E. Brewer
中科院分区:
其他
文献类型:
--
作者:
Mehul A. Shah;J. Hellerstein;E. Brewer

文献摘要

被引文献

相似文献

我们提出了一种屏蔽集群中故障的技术,为长期运行的并行数据流提供高可用性和容错性。我们可以使用这些数据流来实现各种需要高吞吐量、全天候运行的连续查询(CQ)应用程序。例如,网络监控、电话呼叫处理、点击流处理和在线金融分析。我们的主要贡献是一种方案,该方案仔细地将用于分区并行的传统查询处理技术与用于高可用性的进程对方法相结合。这种微妙的集成允许我们在不牺牲结果质量的情况下容忍并行数据流的部分故障。发生故障时,我们的技术提供快速故障转移,并在运行中自动恢复丢失的部件。与在每个数据流基础上直接应用进程对技术相比,这种逐段恢复对正在进行的数据流计算的干扰最小,并提高了可靠性。因此,我们的技术提供了关键CQ应用程序所需的高可用性。我们的技术封装在一个名为Flux的可重用数据流操作符中,它是Exchange的一个扩展,用于组成并行数据流。将容错逻辑封装到Flux中可以最大限度地减少对现有操作符代码的修改,并减轻操作符编写者重复实现和验证这一关键逻辑的负担。我们通过在Telegraph CQ代码库中实现Flux来演示这些特性[8]。
We present a technique that masks failures in a cluster to provide high availability and fault-tolerance for long-running, parallelized dataflows. We can use these dataflows to implement a variety of continuous query (CQ) applications that require high-throughput, 24x7 operation. Examples include network monitoring, phone call processing, click-stream processing, and online financial analysis. Our main contribution is a scheme that carefully integrates traditional query processing techniques for partitioned parallelism with the process-pairs approach for high availability. This delicate integration allows us to tolerate failures of portions of a parallel dataflow without sacrificing result quality. Upon failure, our technique provides quick fail-over, and automatically recovers the lost pieces on the fly. This piecemeal recovery provides minimal disruption to the ongoing dataflow computation and improved reliability as compared to the straight-forward application of the process-pairs technique on a per dataflow basis. Thus, our technique provides the high availability necessary for critical CQ applications. Our techniques are encapsulated in a reusable dataflow operator called Flux, an extension of the Exchange that is used to compose parallel dataflows. Encapsulating the fault-tolerance logic into Flux minimizes modifications to existing operator code and relieves the burden on the operator writer of repeatedly implementing and verifying this critical logic. We present experiments illustrating these features with an implementation of Flux in the TelegraphCQ code base [8].