Overload Management in Data Stream Processing Systems with Latency Guarantees

Overload Management in Data Stream Processing Systems with Latency Guarantees
复制标题

DOI:
--
复制
发表时间:
2012
期刊:
--
影响因子:
--
通讯作者:
Evangelia Kalyviannaki;Themistoklis Charalambous;Marco Fiscato;P. Pietzuch
Evangelia Kalyviannaki;Themistoklis Charalambous;Marco Fiscato;P. Pietzuch
中科院分区:
其他
文献类型:
--
作者:
Evangelia Kalyviannaki;Themistoklis Charalambous;Marco Fiscato;P. Pietzuch

文献摘要

相似文献

流处理系统对于分析由诸如在线社交网络之类的现代应用生成的实时数据正变得越来越重要。它们的主要特点是在实时生成新数据时产生连续的新结果流。流处理系统的资源供应是困难的,由于随时间变化的工作负载数据,引起未知的资源需求。尽管开发了旨在提供工作负载变化的可扩展流处理系统,但仍然存在这样的系统面临瞬时资源短缺的情况。在过载期间,缺乏实时处理所有传入数据的资源;数据在内存中累积,并且其处理延迟不受控制地增长,从而损害流处理结果的新鲜度。在本文中,我们提出了一种反馈控制方法来设计一个非线性离散时间控制器,该控制器不知道要控制的系统或数据的工作负载,并且仍然能够控制单节点流处理系统中的平均元组端到端延迟。结果表明,我们的原型流处理系统的评估,我们的方法控制的平均元组端到端的延迟,尽管随时间变化的工作负载的需求和不断增加的查询数量。
Stream processing systems are becoming increasingly important to analyse real-time data generated by modern applications such as online social networks. Their main characteristic is to produce a continuous stream of fresh results as new data are being generated at real-time. Resource provisioning of stream processing systems is difficult due to time-varying workload data that induce unknown resource demands over time. Despite the development of scalable stream processing systems, which aim to provision for workload variations, there still exist cases where such systems face transient resource shortages. During overload, there is a lack of resources to process all incoming data in real-time; data accumulate in memory and their processing latency grows uncontrollably compromising the freshness of stream processing results. In this paper, we present a feedback control approach to design a nonlinear discrete-time controller that has no knowledge of the system to be controlled or the workload for the data and is still able to control the average tuple end-to-end latency in a single-node stream processing system. The results, of our evaluation on a prototype stream processing system, show that our method controls the average tuple end-to-end latency despite the time-varying workload demands and increasing number of queries.