Data Stream Processing with Concurrency Control

Data Stream Processing with Concurrency Control
复制标题

具有并发控制的数据流处理

DOI:
10.1145/2505420.2505425
复制
发表时间:
2013
期刊:
ACM SIGAPP Applied Computing Review
影响因子:
--
通讯作者:
and Hiroyuki Kitagawa
and Hiroyuki Kitagawa
中科院分区:
--
文献类型:
--
作者:
Masafumi Oyamada;Hideyuki Kawashima;and Hiroyuki Kitagawa

文献摘要

相似文献

数据流处理的最新趋势显示了高级连续查询(CQ)的使用,这些查询引用非流资源,例如数据库和机器学习模型中的关系数据。由于非流式资源可以在多个系统之间共享,因此资源可以在CQ执行期间由系统更新。因此,CQ可能不一致地引用资源,并导致从不适当的结果到致命的系统故障的广泛问题。本文通过将事务处理的概念引入到数据流处理中来解决这一不一致性问题,在本文的第一部分中,我们引入了CQ派生事务的概念,即从CQ派生只读事务,并举例说明了通过保证派生事务和资源更新事务的可串行化来解决不一致性问题。为了保证可串行化,我们提出了三种基于并发控制技术的CQ处理策略:两阶段锁策略、快照策略和乐观策略。实验结果表明,我们提出的CQ处理策略能够保证正确的结果,并且其性能与传统的可能产生错误结果的策略相当。在本文的第二部分中,我们试图从算子调度的角度来提高我们提出的策略的性能。我们注意到我们提出的策略的一个特点:运营商可以重新评估,以防止非串行化的时间表造成性能下降。我们发现的事实,运营商重新评估的数量取决于运营商调度,并提出了一个调度约束,减少重新评估。实验研究表明,我们的约束的有效性:如果我们提出的约束操作员调度,吞吐量增加到5.2倍相比,没有约束的朴素调度。
A recent trend in data stream processing shows the use of advanced continuous queries (CQs) that reference non-streaming resources such as relational data in databases and machine learning models. Since non-streaming resources could be shared among multiple systems, resources may be updated by the systems during the CQ-execution. As a consequence, CQs may reference resources inconsistently, and lead to a wide range of problems from inappropriate results to fatal system failures. In this paper, we address this inconsistency problem by introducing the concept of transaction processing onto data stream processing.In the first part of this paper, we introduceCQ-derived transaction, a concept that derives read-only transactions from CQs, and illustrate that the inconsistency problem is solved by ensuring serializability of derived transactions and resource updating transactions. To ensure serializability, we propose three CQ-processing strategies based on concurrency control techniques: two-phase lock strategy, snapshot strategy, and optimistic strategy. Experimental study shows our CQ-processing strategies guarantee proper results, and their performances are comparable to the performance of conventional strategy that could produce improper results.In the second part of this paper, we try to improve the performance of our proposed strategies from the viewpoint of operator scheduling. We notice a characteristic of our proposed strategies: operators could be re-evaluated to prevent non-serializable schedules causing performance degradation. We find the fact that the number of operator re-evaluation depends on operator scheduling, and propose a scheduling constraint that reduces the re-evaluation. Experimental study shows our constraint's effectiveness: if we add the proposed constraint to operator scheduling, throughput increases up to 5.2 times compared to the naïve scheduling without the constraint.