On Density-Based Data Streams Clustering Algorithms: A Survey

On Density-Based Data Streams Clustering Algorithms: A Survey
复制标题

DOI:
10.1007/s11390-014-1416-y
复制
发表时间:
2014-01-01
影响因子:
1.9
通讯作者:
Saboohi, Hadi
Saboohi, Hadi
中科院分区:
计算机科学3区
文献类型:
--
作者:
Amini, Amineh;Teh, Ying Wah;Saboohi, Hadi

文献摘要

被引文献

相似文献

在过去的几年里,由于数据流的不断增长,聚类数据流引起了人们的广泛关注。数据流对集群提出了额外的挑战,例如有限的时间和内存以及一次性集群。此外,发现具有任意形状的簇在数据流应用中非常重要。数据流是无限的,并随着时间的推移而不断变化,我们对集群的数量没有任何了解。在数据流环境中,由于各种因素,偶尔会出现一些噪声。基于密度的聚类方法是数据流聚类中的一个显著类别,它具有发现任意形状簇和检测噪声的能力。此外,它不需要预先知道簇的数量。由于数据流的特点,传统的基于密度的聚类方法不再适用。近年来,许多基于密度的聚类算法被扩展到数据流。这些算法的主要思想是在聚类过程中使用基于密度的方法,同时克服数据流本身的特性所带来的约束。本文的目的是阐明一些算法在文献中的基于密度的数据流聚类。我们不仅总结了主要的基于密度的数据流聚类算法,讨论他们的独特性和局限性,但也解释了他们如何解决聚类数据流的挑战。此外,我们调查的评价指标,用于验证集群质量和衡量算法的性能。希望这项调查将作为一个踏脚石研究人员研究数据流聚类,特别是基于密度的算法。
Clustering data streams has drawn lots of attention in the last few years due to their ever-growing presence. Data streams put additional challenges on clustering such as limited time and memory and one pass clustering. Furthermore, discovering clusters with arbitrary shapes is very important in data stream applications. Data streams are infinite and evolving over time, and we do not have any knowledge about the number of clusters. In a data stream environment due to various factors, some noise appears occasionally. Density-based method is a remarkable class in clustering data streams, which has the ability to discover arbitrary shape clusters and to detect noise. Furthermore, it does not need the number of clusters in advance. Due to data stream characteristics, the traditional density-based clustering is not applicable. Recently, a lot of density-based clustering algorithms are extended for data streams. The main idea in these algorithms is using density-based methods in the clustering process and at the same time overcoming the constraints, which are put out by data stream's nature. The purpose of this paper is to shed light on some algorithms in the literature on density-based clustering over data streams. We not only summarize the main density-based clustering algorithms on data streams, discuss their uniqueness and limitations, but also explain how they address the challenges in clustering data streams. Moreover, we investigate the evaluation metrics used in validating cluster quality and measuring algorithms' performance. It is hoped that this survey will serve as a steppingstone for researchers studying data streams clustering, particularly density-based algorithms.