Elastic Stream Processing with Latency Guarantees

Elastic Stream Processing with Latency Guarantees
复制标题

DOI:
10.1109/icdcs.2015.48
复制
发表时间:
2015-06
期刊:
2015 IEEE 35th International Conference on Distributed Computing Systems
影响因子:
--
通讯作者:
Björn Lohrmann;P. Janacik;O. Kao
Björn Lohrmann;P. Janacik;O. Kao
中科院分区:
其他
文献类型:
--
作者:
Björn Lohrmann;P. Janacik;O. Kao

文献摘要

被引文献

相似文献

科学和工业中的许多大数据应用已经出现,这些应用需要以低延迟分析大量的流数据或事件数据。本文提出了一种反应式策略,以执行可扩展的流处理引擎(SPE)上运行的数据流的延迟保证,同时最大限度地减少资源消耗。我们引入了一个模型,用于估计延迟的数据流,当内部的任务的并行度发生变化。我们描述了如何持续测量模型的必要性能指标,以及如何通过在运行时确定适当的缩放操作来执行延迟保证。因此,它利用了通用云技术和集群资源管理系统固有的弹性。我们已经实施了我们的战略,作为Nephele SPE的一部分。为了展示我们方法的有效性,我们对一个大型商品集群进行了实验评估,使用合成工作负载以及对真实世界社交媒体数据进行实时情感分析的应用程序。
Many Big Data applications in science and industry have arisen, that require large amounts of streamed or event data to be analyzed with low latency. This paper presents a reactive strategy to enforce latency guarantees in data flows running on scalable Stream Processing Engines (SPEs), while minimizing resource consumption. We introduce a model for estimating the latency of a data flow, when the degrees of parallelism of the tasks within are changed. We describe how to continuously measure the necessary performance metrics for the model, and how it can be used to enforce latency guarantees, by determining appropriate scaling actions at runtime. Therefore, it leverages the elasticity inherent to common cloud technology and cluster resource management systems. We have implemented our strategy as part of the Nephele SPE. To showcase the effectiveness of our approach, we provide an experimental evaluation on a large commodity cluster, using both a synthetic workload as well as an application performing real-time sentiment analysis on real-world social media data.