Extension of DBSCAN in Online Clustering: An Approach Based on Three-Layer Granular Models

Extension of DBSCAN in Online Clustering: An Approach Based on Three-Layer Granular Models
复制标题

DOI:
10.3390/app12199402
复制
发表时间:
2022-09
期刊:
影响因子:
--
通讯作者:
Xinhui Zhang;Xun Shen;Tinghui Ouyang
Xinhui Zhang;Xun Shen;Tinghui Ouyang
中科院分区:
--
文献类型:
--
作者:
Xinhui Zhang;Xun Shen;Tinghui Ouyang

文献摘要

相似文献

在大数据分析中,传统的聚类算法在处理非线性空间数据集时存在精度低、计算成本高等局限性。针对这些问题,本文提出了一种新的在线聚类DBSCAN扩展算法,该算法由三层组成,考虑了DBSCAN、粒度计算(GrC)和基于模糊规则的建模。首先,利用DBSCAN算法在提取结构信息方面的优势,通过DBSCAN将空间数据聚类成结构簇,然后通过GrC由结构信息颗粒(IG)进行描述。其次,基于结构IG,在介质空间中构建一系列粒状模型,并利用它们形成模糊规则来指导空间数据的聚类。最后,借助结构IG和粒度规则,在输出空间中构建基于规则的建模方法以进行在线聚类。本文在合成玩具数据集和典型空间数据集上进行了实验。数值结果验证了该方法在在线空间数据聚类中的可行性。此外,与传统方法和现有 DBSCAN 变体的比较研究证明了该方法的优越性,以及精度的提高和计算开销的减少。
In big data analysis, conventional clustering algorithms have limitations to deal with nonlinear spatial datasets, e.g., low accuracy and high computation cost. Aiming at these problems, this paper proposed a new DBSCAN extension algorithm for online clustering, which consists of three layers, considering DBSCAN, granular computing (GrC), and fuzzy rule-based modeling. Firstly, making use of DBSCAN algorithms’ advantages at extracting structural information, spatial data are clustered via DBSCAN into structural clusters, which are subsequently described by structural information granules (IG) via GrC. Secondly, based on the structural IGs, a series of granular models are constructed in the medium space, and utilized to form fuzzy rules to guide clustering on spatial data. Finally, with the help of structural IGs and granular rules, a rule-based modeling method is constructed in the output space for online clustering. Experiments on a synthetic toy dataset and a typical spatial dataset are implemented in this paper. Numerical results validate the feasibility to the proposed method in online spatial data clustering. Moreover, comparative studies with conventional methods and existing DBSCAN variants demonstrate the superiorities of the proposed method, as well as accuracy improvement and computation overhead reduction.