Optimizing Timeliness and Cost in Geo-Distributed Streaming Analytics

Optimizing Timeliness and Cost in Geo-Distributed Streaming Analytics
复制标题

优化地理分布式流分析的及时性和成本

DOI:
10.1109/tcc.2017.2750678
复制
发表时间:
2020
影响因子:
6.5
通讯作者:
R. Sitaraman
R. Sitaraman
中科院分区:
计算机科学2区
文献类型:
--
作者:
B. Heintz;A. Chandra;R. Sitaraman

文献摘要

被引文献

相似文献

快速的数据流不断地从不同的来源产生,包括位于全球各地的用户、设备和传感器。这就需要高效的地理分布式流分析来提取及时的信息。典型的地理分布式分析服务使用轮辐模型,包括由广域网(WAN)连接到中央数据仓库的多个边缘。在本文中,我们重点讨论了广泛使用的窗口分组聚合原语,并研究了在边缘和中心应该执行多少计算的问题。我们开发了算法来优化两个关键指标:广域网流量和延迟(获得结果的延迟)。我们提出了一系列最优的离线算法,它们共同最小化这些指标,并且我们使用这些算法来指导我们设计实用的在线算法,这些算法基于这样的见解,即窗口分组聚合可以建模为缓存大小随时间变化的缓存问题。我们通过部署在PlanetLab上的Apache Storm实现来评估我们的算法。使用来自大型商业CDN的流行分析服务的匿名跟踪的工作负载,我们的实验表明,我们的在线算法对于各种系统配置,流到达率和查询实现了近乎最佳的流量和过时性。
Rapid data streams are generated continuously from diverse sources including users, devices, and sensors located around the globe. This results in the need for efficient geo-distributed streaming analytics to extract timely information. A typical geo-distributed analytics service uses a hub-and-spoke model, comprising multiple edges connected by a wide-area-network (WAN) to a central data warehouse. In this paper, we focus on the widely used primitive of windowed grouped aggregation, and examine the question of how much computation should be performed at the edges versus the center. We develop algorithms to optimize two key metrics: WAN traffic and staleness(delay in getting results). We present a family of optimal offline algorithms that jointly minimize these metrics, and we use these to guide our design of practical online algorithms based on the insight that windowed grouped aggregation can be modeled as a caching problem where the cache size varies over time. We evaluate our algorithms through an implementation in Apache Storm deployed on PlanetLab. Using workloads derived from anonymized traces of a popular analytics service from a large commercial CDN, our experiments show that our online algorithms achieve near-optimal traffic and staleness for a variety of system configurations, stream arrival rates, and queries.