An incremental clustering method based on the boundary profile.

An incremental clustering method based on the boundary profile.
复制标题

DOI:
10.1371/journal.pone.0196108
复制
发表时间:
2018
期刊:
影响因子:
3.7
通讯作者:
Wu G
Wu G
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Bao J;Wang W;Yang T;Wu G

文献摘要

参考文献

被引文献

相似文献

许多重要的应用程序不断产生数据,如金融事务管理、卫星监控、网络流量监控和web信息处理。数据挖掘的结果总是随着新生成的数据而变化。显然,对于聚类任务,最好是基于旧数据增量地更新新的聚类结果,而不是从头开始对所有数据重新聚类。增量聚类方法是解决大数据增长型聚类问题的重要途径。提出了一种基于边界轮廓的增量聚类方法,用于在动态增长的数据集上寻找任意形状的聚类。该方法将现有的聚类结果用一组边界轮廓来表示,而不是保留所有的数据,而是丢弃聚类的内部点。大大节省了时间和空间的存储成本。为了识别边界轮廓,本文提出了一种基于边界向量的边界点检测(BV-BPD)算法,该算法总结了现有聚类的结构。BPIC方法以在线方式处理每个新点,并以批处理方式更新聚类结果。当一个新的点到达时,BPIC方法根据新数据和边界轮廓之间的关系,要么立即标记它,要么临时将它放入桶中。该方法利用桶来区分噪声和新聚类的潜在种子,减轻数据顺序的影响。当桶已满时,BPIC方法将对其中的数据进行聚类并更新聚类结果。因此,BPIC方法对噪声和新数据的顺序不敏感,这对增量聚类过程的鲁棒性至关重要。在实验中,对边界点检测算法BV-BPD的性能与现有方法进行了比较。结果表明,BV-BPD方法优于现有方法。此外,从聚类质量、时间和空间效率等方面考察了BPIC和其他两种增量聚类方法的性能。实验结果表明,BPIC方法能够在大数据集上获得合格的聚类结果,具有较高的时间和空间效率。
Many important applications continuously generate data, such as financial transaction administration, satellite monitoring, network flow monitoring, and web information processing. The data mining results are always evolving with the newly generated data. Obviously, for the clustering task, it is better to incrementally update the new clustering results based on the old data rather than to recluster all of the data from scratch. The incremental clustering approach is an essential way to solve the problem of clustering with growing Big Data. This paper proposes a boundary-profile-based incremental clustering (BPIC) method to find arbitrarily shaped clusters with dynamically growing datasets. This method represents the existing clustering results with a collection of boundary profiles and discards the inner points of clusters rather than keep all data. It greatly saves both time and space storage costs. To identify the boundary profile, this paper presents a boundary-vector-based boundary point detection (BV-BPD) algorithm that summarizes the structure of the existing clusters. The BPIC method processes each new point in an online fashion and updates the clustering results in a batch mode. When a new point arrives, the BPIC method either immediately labels it or temporarily puts it into a bucket according to the relationship between the new data and the boundary profiles. A bucket is employed to distinguish the noise from the potential seeds of new clusters and alleviate the effects of data order. When the bucket is full, the BPIC method will cluster the data within it and update the clustering results. Thus, the BPIC method is insensitive to noise and the order of new data, which is critical for the robustness of the incremental clustering process. In the experiments, the performance of the boundary point detection algorithm BV-BPD is compared with the state-of-the-art method. The results show that the BV-BPD is better than the state-of-the-art method. Additionally, the performance of BPIC and other two incremental clustering methods are investigated in terms of clustering quality, time and space efficiency. The experimental results indicate that the BPIC method is able to get a qualified clustering result on a large dataset with higher time and space efficiency.
基于三支决策理论的基于树的增量重叠聚类方法
DOI: 10.1016/j.knosys.2015.05.028
发表时间: 2016-01-01
影响因子: 8.8
作者:
Yu, Hong;Zhang, Cong;Wang, Guoyin
通讯作者: Wang, Guoyin
DOI: 10.3390/a5030364
发表时间: 2012-09-01
期刊: ALGORITHMS
影响因子: 2.3
作者:
Azzopardi, Joel;Staff, Christopher
通讯作者: Staff, Christopher
DOI: 10.1007/s11390-014-1416-y
发表时间: 2014-01-01
影响因子: 1.9
作者:
Amini, Amineh;Teh, Ying Wah;Saboohi, Hadi
通讯作者: Saboohi, Hadi
DOI: 10.1016/j.patcog.2005.01.025
发表时间: 2005-11-01
影响因子: 8
作者:
Liao, TW
通讯作者: Liao, TW
DOI: 10.1109/tkde.2006.38
发表时间: 2006-03-01
影响因子: 8.9
作者:
Xia, CY;Hsu, W;Ooi, BC
通讯作者: Ooi, BC