DistStream: An Order-Aware Distributed Framework for Online-Offline Stream Clustering Algorithms

DistStream: An Order-Aware Distributed Framework for Online-Offline Stream Clustering Algorithms
复制标题

DOI:
10.1109/icdcs47774.2020.00075
复制
发表时间:
2020-11
期刊:
2020 IEEE 40th International Conference on Distributed Computing Systems (ICDCS)
影响因子:
--
通讯作者:
Lijie Xu;Xingtong Ye;Kai Kang;Tian Guo;Wensheng Dou;Wei Wang-;Jun Wei
Lijie Xu;Xingtong Ye;Kai Kang;Tian Guo;Wensheng Dou;Wei Wang-;Jun Wei
中科院分区:
其他
文献类型:
--
作者:
Lijie Xu;Xingtong Ye;Kai Kang;Tian Guo;Wensheng Dou;Wei Wang-;Jun Wei

文献摘要

相似文献

流聚类是一种重要的数据挖掘技术,用于捕捉实时数据流中的演变模式。如今的数据流,例如物联网事件和网络点击,通常速度很快且包含动态变化的模式。现有的流聚类算法通常遵循在线 - 离线范式以及一次更新一条记录的模型,这种模型是为在单机上运行而设计的。这些具有顺序更新模型的流聚类算法无法高效地并行化,并且无法为流聚类提供所需的高吞吐量。在本文中,我们提出了DistStream,这是一个能够有效扩展在线 - 离线流聚类算法的分布式框架。为了使这些算法并行化以实现高吞吐量,我们开发了一种具有高效并行化方法的小批量更新模型。为了保持较高的聚类质量,DistStream的小批量更新模型在并行执行期间的所有计算步骤中保留更新顺序,这能够反映动态变化的流数据的最新变化。我们在Spark Streaming之上实现了DistStream,以及基于DistStream的四种具有代表性的流聚类算法。我们对三个真实世界数据集的评估表明,基于DistStream的流聚类算法能够实现次线性的吞吐量增益,并且与它们的单机对应算法具有相当(99%)的聚类质量。
Stream clustering is an important data mining technique to capture the evolving patterns in real-time data streams. Today’s data streams, e.g., IoT events and Web clicks, are usually high-speed and contain dynamically-changing patterns. Existing stream clustering algorithms usually follow an online-offline paradigm with a one-record-at-a-time update model, which was designed for running in a single machine. These stream clustering algorithms, with this sequential update model, cannot be efficiently parallelized and fail to deliver the required high throughput for stream clustering.In this paper, we present DistStream, a distributed framework that can effectively scale out online-offline stream clustering algorithms. To parallelize these algorithms for high throughput, we develop a mini-batch update model with efficient parallelization approaches. To maintain high clustering quality, DistStream’s mini-batch update model preserves the update order in all the computation steps during parallel execution, which can reflect the recent changes for dynamically-changing streaming data. We implement DistStream atop Spark Streaming, as well as four representative stream clustering algorithms based on DistStream. Our evaluation on three real-world datasets shows that DistStream-based stream clustering algorithms can achieve sublinear throughput gain and comparable (99%) clustering quality with their single-machine counterparts.